Earlier quoted context omitted.
Hasn't this always been the case? Arxiv being used for self promotion and Kaggle being used to pivot into the industry. It is not a recent phenomenon.
I cannot speak for the intended purpose of ArXiv by its creators but I can tell you that, in the conference circles, its main intended use was flag planting: people were afraid that their competitors would tweet some results while their (earlier) paper was under anonymous review, and so researchers started putting their stuff on ArXiv first to ensure no one would "steal" their claim of being there first.
Evidence of inconsistencies in evaluation process and selection of winners
281–290 of 338 posts
Re: Evidence of inconsistencies in evaluation process and selection of winners
#282Re: Evidence of inconsistencies in evaluation process and selection of winners
#283Re: Evidence of inconsistencies in evaluation process and selection of winners
#284Earlier quoted context omitted.
One of my first gigs as a consultant was to write a project management system for a company that didn't really need a custom project management system. The CEO pulled me aside and told me the only important feature of the project management system was that you couldn't assign the same priority to two features. I would be blamed for making such a crappy project management system, but that's what I was there for. Once…
> the only important feature of the project management system was that you couldn't assign the same priority to two feature That is a good idea for a project management system. Force ranking of priorities.
Re: Evidence of inconsistencies in evaluation process and selection of winners
#285First off, I want to set some context on the AGI hackathon. This was co-organized by Kaggle and Google DeepMind, and we had ~20 judges from both organizations. The hackathon concluded on Apr 16 and we had initially anticipated a judging period of 1.5 months (till May 31). However, we ended up extending the judging period by another 1.5 months (to Jul 13) because we wanted to do right by participants.
Second, I want to emphasize and unequivocally clarify that every single winning submission went through at least 2 human judges, and in some cases, up to 3-4 human judges. These judges reviewed and scored the submissions independently based on the rubric we highlighted on the hackathon page.
Thirdly, I acknowledge that there is always an element of human subjectivity to reviewing qualitative submissions in hackathons. As best we can, we have put in place processes that ensure rigorous human review against objective standards and to reduce the possibility of bias by having multiple independent judges. We understand there may be valid disagreement over outcomes, but hopefully the above context clarifies this was not carelessly outsourced to LLM judges.
Thanks, Nick
Re: Evidence of inconsistencies in evaluation process and selection of winners
#286Hi all, I'm Nick, Product Manager for Kaggle Benchmarks and one of the co-organizers and judges for this AGI hackathon. First off, I want to set some context on the AGI hackathon. This was co-organized by Kaggle and Google DeepMind, and we had ~20 judges from both organizations. The hackathon concluded on Apr 16 and we had initially anticipated a judging period of 1.5 months (till May 31). However, we ended up extend…
Re: Evidence of inconsistencies in evaluation process and selection of winners
#287Earlier quoted context omitted.
With the exception of _one_ company that I worked at, pretty much every[0] company was a struggle between engineering and management. Engineering wants to get the software correct, and management wants to fire-hose features into the market. Most of the time (so more than half, at least), management tends to have a compulsion to mindlessly imitate what other companies/competitors are doing, usually without prioritizat…
To be the devil's advocate, engineers themselves can be hilariously incompetent too. Engineers have a tendency of assuming that budget is infinite and target audience is other engineers from same specialization. Open-source projects often have this problem where you can have dozens of thousands of man-hours poured into a project without a single end-user opinion taken into account. At some point my manager, who himse…
How is this "a problem"? The reason there are dozens of thousands of hours on an open source project is because the end-users are working on it. Some projects exist solely for someone to work on it (that is, the "working on it" is the "use case"). Open source does not expect or need to make money or get "users", so how people discover or source what they build and how to proceed isn't really a "problem".
Re: Evidence of inconsistencies in evaluation process and selection of winners
#288Hi all, I'm Nick, Product Manager for Kaggle Benchmarks and one of the co-organizers and judges for this AGI hackathon. First off, I want to set some context on the AGI hackathon. This was co-organized by Kaggle and Google DeepMind, and we had ~20 judges from both organizations. The hackathon concluded on Apr 16 and we had initially anticipated a judging period of 1.5 months (till May 31). However, we ended up extend…
Why not address the objective evidence the OP provided? To an impartial observer, it seemed quite overwhelming.
Writeup quality is 20% of the total grade and there are other factors like dataset quality and results, which hold a much larger weight to the scores.
This has been made public to participants since day 1 of the launch and we've adhered very closely to this rubric:)
Re: Evidence of inconsistencies in evaluation process and selection of winners
#289Earlier quoted context omitted.
Why not address the objective evidence the OP provided? To an impartial observer, it seemed quite overwhelming.
I don't want to comment on and single out any one participant, but please check out the rubric on how we did the grading: https://www.kaggle.com/competitions/kaggle-measuring-agi/ove... Writeup quality is 20% of the total grade and there are other factors like dataset quality and results, which hold a much larger weight to the scores. This has been made public to participants since day 1 of the launch and we've adher…
So it must have been the "results" that moved the needle :D
Re: Evidence of inconsistencies in evaluation process and selection of winners
#290Earlier quoted context omitted.
That just isn’t true. AI is capable of performing a lot of grunt work reliably. Still must be reviewed. But a big productivity gain over doing everything yourself.
While I agree with you in principle, I think the parent has a point here: where's the amazing product that couldn't have been done without AI? By now we should have seen some major new invention/company, incredibly fast revolutionary feature rollouts etc but I'm just seeing more of the same.
My phone can search my photos with text based on what's represented in the photo. That wasn't a thing for computers from 1950 through about 2012, and now it's just normal. Google Lens will tell you what a plant is, Merlin app will identify birds from birdsong.
In 2000 we had Dragon Naturally Speaking and Kurzweil VoicePad, now we have voice recognition and machine transcription that works usefully well. 'Now' being sometime after 2012.
I don't remember what we had for human language translation, but since transformers and LLMs and GP-GPU, we have Google Translate and a lot of competitors, such that usefully good translation is everywhere - in the right click menu in FireFox, for example. Google translate will take a photo with foreign text in it, translate the text, overlay it on the photo.
Useful programming language porting, it's gone from basically impossible or unusable, to very usable.
Text synthesis, my employer has an AI generated newsletter. It's not something I want, but it's also not something you could do at all ten years ago.
Script synthesis, it's now easy to one-shot a few lines of script or command line to do something, faster than opening the ffmpeg man pages or whatever. "that didn't work, here's the error: " "the codec is not supported in this mode, try --work-please" and
Image synthesis. My local daily newspaper has at least one full page AI image, that's not something you could do at all ten years ago.
Camera drones, now they can clock onto your face, fly while following you and keeping the you in shot, and respond to gestures for things like landing.
> "By now we should have seen some major new invention"
There's no law of the Universe which says this, "you should see a new invention k years after the invention of a transformer architecture". Solar electricity generation was first noticed in 1839, first commercialised in 1954, and it's only just (give or take a few years) become a significant percentage of world power generation (6% in 2022) and the cheapest form of electricity.
"AI" was considered a hundred years ago with 'Robots' and Turing Tests, wasn't usefully progressed until computers in say the 1950's with LISP and chess, the 1970s with Prolog, but had a huge shift around 2012 with GP-GPU and powerful enough video cards to crunch large amounts of data and with the internet providing large amounts of data, then again with transformer architectures and LLMs, and then with tensor flow and TPU/NPU dedicated accelerators.
Do you expect the world will look the same in 2036? 2046? That all this investment in AI datacenters will just stop and go away? That all the dedicated cores in mobile devices will fall unused? That Waze and Tesla Full Self Driving will just stop where they are even if processing power 10x's? That Amazon will stop rolling out warehouse robots because humans are better? That drones and robo-mowers, robo vacuums, Alexas and OK Google's are going to be no different? That nobody will train anything else, connect up any other inputs or outputs, try any new uses, or make use of any changing costs?