SYNTHESIS NOTE
TopicsEvaluationsthis note

Can fairness frameworks extend to general-purpose language models?

Existing fairness frameworks were designed for narrow, structured tasks. This explores whether they scale to LLMs, which serve multiple populations, sensitive attributes, and use cases simultaneously.

Synthesis note · 2026-06-03 · sourced from Evaluations

Machine-learning fairness frameworks — group fairness, fair representations — were built for systems with well-structured inputs and outputs and a self-evident use (lending, recidivism, coreference). This work argues they break on general-purpose LLMs: each framework either does not logically extend to LLM tasks (unstructured natural-language data) or becomes intractable, because LLMs touch a multitude of populations, sensitive attributes, and use cases at once. The conclusion is sharp: it is not feasible to certify or guarantee that an LLM is generally "fair."

The constructive move is to lower the target from universal fairness to use-case-specific fairness, governed by three guidelines: the criticality of context, the responsibility of LLM developers, and stakeholder participation in an iterative design-and-evaluation process — with the speculative note that AI's general capabilities might eventually help address fairness as a form of scalable AI-assisted alignment.

The keeper is the impossibility-style framing: "fair LLM" is not a global property you can stamp on a model; fairness only has teeth relative to a context and its stakeholders.

This pairs with the vault's preference/alignment-pluralism thread. It mirrors Can human-centered LLM design ever achieve universal solutions? and complements Can a single reward model represent diverse human preferences? — both reject one-size-fits-all normative targets — and supports Should AI alignment target preferences or social role norms? as the contextual alternative.

Inquiring lines that read this note 4

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why can LLMs generate ideas better than they evaluate them? How can AI alignment serve diverse human preferences at scale? What makes specific clarifying questions more effective than generic ones?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 100 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

group-fairness frameworks do not extend to general-purpose LLMs so fairness must be pursued per use-case