DeepMind Kaggle Winners: New Benchmarks Reveal True AI Reasoning Beyond Recall

Blatant AI slop just won a 25k USD DeepMind Kaggle Grand Prize

DeepMind Kaggle Winners: New Benchmarks Reveal True AI Reasoning Beyond Recall

I am thrilled to announce the winners of the DeepMind Kaggle hackathon, where over 1,000 teams created benchmarks to measure genuine AI reasoning. The grand prize winners developed tools like MEDLEY-BENCH and LearningBench to test how models handle social pressure, learn new rules instantly, and recognize their own uncertainty. These innovations move us beyond simple data recall toward evaluating true cognitive abilities in frontier models.

A model that reports 95% confidence and never abstains is more dangerous than one that gets questions wrong, because it gives the human operator no signal that something might be off.
  1. ecshafer

    AI is useful. But the amount of people that are simply offloading all of their thinking to AI and blindly accepting the answer is absurd. Kaggle is most likely using ai to assess the submissions and are not using any common sense by blindly accepting the results.

  2. hoppp

    I don't know about this exact competition but overall fair hackathons have been killed by AI.

    It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners.

    It used to be about human skill, now it's about ideas and of course insiders are the main winners.

  3. leo3191

    Hi all, I'm Nick, Product Manager for Kaggle Benchmarks and one of the co-organizers and judges for this AGI hackathon.

    First off, I want to set some context on the AGI hackathon. This was co-organized by Kaggle and Google DeepMind, and we had ~20 judges from both organizations. The hackathon concluded on Apr 16 and we had initially anticipated a judging period of 1.5 months (till May 31). However, we ended up extending the judging period by another 1.5 months (to Jul 13) because we wanted to do right by participants.

    Second, I want to emphasize and unequivocally clarify that every single winning submission went through at least 2 human judges, and in some cases, up to 3-4 human judges. These judges reviewed and scored the submissions independently based on the rubric we highlighted on the hackathon page.

    Thirdly, I acknowledge that there is always an element of human subjectivity to reviewing qualitative submissions in hackathons. As best we can, we have put in place processes that ensure rigorous human review against objective standards and to reduce the possibility of bias by having multiple independent judges. We understand there may be valid disagreement over outcomes, but hopefully the above context clarifies this was not carelessly outsourced to LLM judges.

    Thanks,

    Nick

  4. throwfaraway135

    AI submissions and AI judges a match made in (AI) heaven.

  5. mr_toad

    People have been using brute force methods to win Kaggle competitions since the beginning (and other people have been complaining about it just as long).

    At its core, ML is all about computer generated models (automated feature selection, hyper parameter tuning). Many (most?) of the models produced in Kaggle are already black boxes and have been for a long time. The model that won the Netflix prize was never used in production for that reason.

    Using an LLM to generate code to generate a black box is pretty much par for the course.

  6. ryukoposting

    I thought Kaggle was a website where you download dubious CSV files of annualized bean consumption in Bolivia, or whatever.

    Was Kaggle ever a reputable source of original research, or a source of anything with any provenance at all? That would be news to me. The fact that 25 grand was involved this time is unique, I guess.

  7. apwheele

    I think this is a good meta-lesson for Kaggle. When you have objective metrics to hill-climb towards, AI can do quite well. When you just phone it in and rely on LLM as a Judge, the results are not so great.

  8. sublimefire

    Went through the comments here and there and one thing to note is that there was a question about who do you think should have won instead. This is a good question because it is possible that all submissions were like this or there were ones that looked just worse. It would be quite useful to know who came close as well in this case. If you knew which submissions were good you could have a process to revoke the prize and give it to someone else in case of fraud or negligence or similar.

    Having said that it is also possible that the mistakes and claims were a human error, sure a lot gets ai generated these days but there is a chance in which case the accusation does not look so severe anymore.

More from this day

2026-07-17