INQUIRING LINE

When a platform retires a competition format, should its old medals keep full weight, or be discounted or labeled with context?

Should platforms downweight or relabel credentials after retiring the format that issued them?

This explores what a platform should do with badges or medals it already handed out once it shuts down the competition format that produced them: quietly lower their weight, add context to them, or leave them as they are.


This explores what should happen to old credentials once the system that issued them is gone. The clearest evidence in the corpus is an audit of Kaggle competition medals. Medals predicted how well someone would perform on hidden tests, but almost all of that predictive power came in the medal's first year. That held both before and after generative AI arrived, so fresh medals stayed meaningful Do Kaggle medals still predict performance after AI arrived?. The problem was old medals. When the platform retired its upload-style competition format, those medals kept losing relevance on the usual schedule but stayed on profiles at full face value. The audit traces about half of the drop in how informative those medals were to this 'stranding', not to AI How much did retiring a competition format hurt medal credibility?. So the damage didn't come from cheating or from the medals being wrong. It came from the platform letting a credential outlive the process that once checked it.

That points toward relabeling over silent downweighting. Since a credential's information fades with age in a predictable way, the honest fix is to show its age and where it came from, for example 'earned in a retired format, 2019', rather than quietly adjusting a hidden score. A related idea from a different area treats distilled expertise as versioned files that can be inspected, corrected and rolled back, not as hidden state Can person-grounded skills remain auditable without hidden prompt state?. Credentials could follow the same rule: a visible lifecycle beats an invisible reweighting that nobody can audit or dispute.

The reason to act at all is the people and systems reading the credentials. They mostly don't look past the label. LLM judges reliably give higher scores to responses that carry authority signals like references or polished formatting, whether or not the content is any good, and these biases can be exploited with no technical skill Can LLM judges be fooled by fake credentials and formatting? Can LLM judges be tricked without accessing their internals?. AI-assisted hiring adds another twist: models prefer resumes they rewrote themselves, which shows evaluators reacting to surface style rather than substance Do language models favor resumes they rewrote themselves?. On the human side, 'cognitive surrender' describes people accepting a signal without checking whether anything still backs it, because checking is costly When do users stop checking whether AI output is actually backed?. A stranded medal is exactly that kind of unbacked signal. Recruiters and automated screeners will keep treating it as current unless the label itself says otherwise.

Here's the less obvious takeaway. A credential's worth depends on something still standing behind it, not just on the fact that it was once earned. Proposals for 'personhood credentials' make this explicit: their value comes from a trusted issuer that stays accountable for what the credential claims Can people prove they are human without revealing who they are?. When a platform retires a format, it withdraws that backing without saying so. Relabeling is how it admits the change.

The corpus has limits here. It gives strong evidence that stranding causes harm and that evaluators trust labels at face value. It has no studies testing whether relabeling or downweighting actually changes how recruiters or algorithms behave. The case for relabeling is an inference from the evidence, not a measured result.


Sources 8 notes

Do Kaggle medals still predict performance after AI arrived?

Across 444,698 participations, medals predicted hidden-test performance almost entirely through their first year in both pre- and post-AI eras. Fresh medals retained most value after generative AI arrived, suggesting verified credentials stayed informative despite platform changes.

How much did retiring a competition format hurt medal credibility?

The audit attributes roughly half the decline in upload-format medal informativeness to institutional stranding: the platform retired the format before AI, medals aged on schedule, yet stayed visible at their original value. This decoupled the credential from the validation mechanism it once represented.

Can person-grounded skills remain auditable without hidden prompt state?

COLLEAGUE.SKILL treats distilled expertise as versioned files subject to inspection, correction, and rollback—not hidden prompt state. Separating capability tracks from behavior tracks enables independent audit of what someone knows versus how they act.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Show all 8 sources
Do language models favor resumes they rewrote themselves?

Across a controlled experiment on 2,245 resumes, eight of nine LLMs preferred their own rewrites over matched human versions when evaluating candidates, with preference rates ranging from 26% to 98%. The bias strengthened in larger models and emerged from stylistic alignment rather than content quality differences.

When do users stop checking whether AI output is actually backed?

Users systematically accept AI outputs without verification because checking is costly and fluent output builds false confidence. This receiver-side surrender—measured in studies showing 80% unchallenged adoption—is what enables inflationary token systems to function at scale.

Can people prove they are human without revealing who they are?

Personhood credentials—privacy-preserving digital credentials issued by trusted institutions—let users prove they are real people rather than AI without revealing personal information. They address three harms: sockpuppets, bot attacks, and misleading agents.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.