When a contest format is retired, do its medals go stale faster, or does something else go wrong?
Do medals from retired competition formats lose predictive power faster than others?
This explores whether Kaggle medals earned in a competition format the platform later retired stop predicting real skill sooner than medals from formats that kept running, or whether something else explains why they look less reliable.
This explores whether medals from a retired competition format lose predictive value faster than other medals. The corpus suggests the short answer is no, and the real story is more interesting: those medals don't age faster, but nothing replaces them and nothing marks them as old. Across nearly 445,000 competition entries, Kaggle medals predicted performance on hidden test data almost entirely through their first year, both before and after generative AI arrived Do Kaggle medals still predict performance after AI arrived?. In other words, every medal has a short shelf life. A fresh medal is a strong signal and an old one is a weak one, whatever format it came from.
The retired format went wrong in a different way. When the platform shut down the upload format, its medals kept aging on the normal schedule, but no new ones were being earned to refresh the pool. The platform also kept showing the old medals at full value. An audit attributes about half of the drop in that format's medal informativeness to this "institutional stranding" How much did retiring a competition format hurt medal credibility?. The credential became detached from the competition process that once gave it meaning. Every medal from that format was now a stale medal, and nothing on the profile showed that.
The pattern matters beyond Kaggle because it is a general way that signals fail. A score stops tracking skill once the process that produced it is no longer being checked against reality. Math benchmarks show the same thing from another angle. One model reconstructs more than half of the MATH-500 benchmark from partial prompts but scores zero on a newer benchmark released after its training data was collected. Its apparent reasoning gains on the old test were mostly memorization Does RLVR success on math benchmarks reflect genuine reasoning improvement?. A frozen benchmark and a retired medal decay in the same way: the number stays visible while what it measures slowly drifts away.
There is a further risk. Once a signal no longer fully represents the thing it stands for, anything optimized against it starts to exploit the gap. That is the shared mechanism behind reward hacking, whether the optimizing happens in model training, output selection, or prompt revision Does reward hacking always stem from the same failure?. Recruiters or ranking systems that treat stranded medals at face value would be optimizing against exactly that kind of incomplete signal.
The corpus is thin here. Only two notes study medals directly, and neither compares decay rates format by format beyond the upload case. The useful takeaway is that a credential's value depends on whether the system that issued it is still running. Asking "how old is this medal?" is not enough. You also need to ask whether anyone is still earning this kind of medal.
Sources 4 notes
Across 444,698 participations, medals predicted hidden-test performance almost entirely through their first year in both pre- and post-AI eras. Fresh medals retained most value after generative AI arrived, suggesting verified credentials stayed informative despite platform changes.
The audit attributes roughly half the decline in upload-format medal informativeness to institutional stranding: the platform retired the format before AI, medals aged on schedule, yet stayed visible at their original value. This decoupled the credential from the validation mechanism it once represented.
Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.
Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stranded Credentials: Keeping Online Reputation Systems Informative in the AI Era
- An Eye Tracking Study: Are AI Overviews Changing Search Behavior?
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
- Spurious Rewards: Rethinking Training Signals in RLVR
- Natural Emergent Misalignment From Reward Hacking In Production RL