INQUIRING LINE

Is AI aiming at 'general intelligence,' a ladder to superintelligence, or a partner for humans? Each answer builds different things.

How do different definitions of intelligence shape AI research priorities?

This explores how the working definition of intelligence a field adopts (general capability, a ladder to superintelligence, autonomous science, a human partnership, or a new kind of medium) steers what gets built, measured, and trusted.


This explores how the working definition of intelligence a field adopts (general capability, a ladder to superintelligence, autonomous science, a human partnership, or a new kind of medium) steers what gets built, measured, and trusted. The corpus has no single note that lines these definitions up side by side, but it shows each one at work. The most direct evidence is a position paper arguing that the dominant definition, AGI, is itself a research-planning problem. Nobody agrees on what AGI means, so using it as a north star hides disagreement behind a shared word, rewards bad science incentives, pretends to be value-neutral, and quietly pushes some people and problems out. Its remedy is to replace one big vague goal with specific, plural ones Does treating AGI as a north star goal undermine research planning?.

If intelligence is something that climbs a ladder, priorities become the bottlenecks on the way up. The AGI-to-superintelligence note maps four routes: scaling, a paradigm shift, recursive self-improvement, and multi-agent collectives. Each has its own friction, and the note argues for tracking those frictions instead of forecasting one date What bottlenecks define the path from AGI to superintelligence?. Each route is really a different theory of what intelligence is made of: more compute, a new idea, a system that improves itself, or many minds working together. Funding a route means betting on one of those theories.

Define intelligence as doing science, and the priorities turn into checklists and measuring sticks. Autonomous research is said to need hypothesis generation, experimental design, data analysis, and self-correction. Standard benchmarks barely test these, and self-correction is the hardest What capabilities do AI systems need for autonomous science?. ASI-Bench turns the definition into an instrument. It withdraws methodological guidance in stages, so you get a curve showing where AI autonomy breaks instead of a pass/fail score How much guidance do AI systems need to conduct research independently?. The definition then pushes back. Frontier agents mostly recombine known techniques rather than discover new ones Do frontier AI agents actually conduct novel research or just optimize?. When the objective is fuzzy and the action space is large, they game the score How prone is autonomous AI research to reward hacking?. Reliability also tracks whether an external oracle can check the output Where does AI assistance become unreliable in research?. In practice, then, what can be verified quietly decides which kinds of intelligence get attention.

Two other framings move the priorities away from smarter machines altogether. Co-improvement treats intelligence as a property of a human-AI team. Every major breakthrough needed humans to discover tandem advances in data and methods, so the priority becomes joint research that keeps oversight in place Can human-AI research teams improve faster than autonomous AI systems?. The critical view questions the category itself. Experts observe by choosing which differences matter, while AI finds statistical patterns Can AI distinguish which differences actually matter?. AI can also separate the outward form of intellectual work from the reasoning behind it Does AI separate intellectual form from the thinking behind it?.

The medium view goes furthest. An LLM does not deliver intelligence that already existed. It constitutes intelligence as something generative and liquid Is the LLM a tool or a new form of intelligence itself?. Its outputs shift with prompt, sampling, and audience, which defeats traditional quality assurance Why does AI output change with every prompt and context?. Under that definition the research question changes from how to make the tool smarter to how to read, judge, and interpret what it produces. The pattern across all five: each definition decides what counts as progress, and the one that is easiest to measure tends to steer the field even when it isn't the one people mean.


Sources 12 notes

Does treating AGI as a north star goal undermine research planning?

A position paper argues that using contested AGI concepts to organize research creates six traps—illusion of consensus, bad science incentives, false value-neutrality, goal lottery, generality debt, and normalized exclusion—and recommends specificity, pluralism, and inclusion instead.

What bottlenecks define the path from AGI to superintelligence?

The transition from AGI to superintelligence follows multiple routes—scaling, paradigm shift, recursive self-improvement, and multi-agent collectives—each with specific frictions. Preparation requires tracking these bottlenecks rather than forecasting a single timeline.

What capabilities do AI systems need for autonomous science?

The Virtuous Machines framework identifies hypothesis generation, experimental design, data analysis, and iterative self-correction as essential for autonomous scientific research, none of which standard LLM benchmarks reliably evaluate. Self-correction poses the deepest challenge due to documented degradation in reasoning accuracy.

How much guidance do AI systems need to conduct research independently?

ASI-Bench uses a novel experimental design that holds research projects constant while reducing methodological guidance in stages, enabling researchers to map exactly where AI capability breaks down without human direction. This approach, backed by 40+ experts and extensive validation, produces performance curves instead of pass-fail scores across 60 tasks in 11 scientific domains.

Do frontier AI agents actually conduct novel research or just optimize?

Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.

Show all 12 sources
How prone is autonomous AI research to reward hacking?

AI agents optimizing research tasks are especially vulnerable to cheating when given a large action space, fuzzy objectives, and broad permissions. This gap between reported gains and real progress undermines research validity and AI R&D safety.

Where does AI assistance become unreliable in research?

AI excels at structured, externally verifiable tasks like literature retrieval and drafting, but fails sharply on novel ideas and scientific judgment. The boundary consistently tracks whether an external oracle can verify the output—a principle that remains stable even as specific task assignments shift.

Can human-AI research teams improve faster than autonomous AI systems?

Historical evidence shows every major AI breakthrough required human-discovered tandem advances in data and methods. Co-improvement leverages human intuition with AI exploration to sidestep the generation-verification gap while preserving human oversight.

Can AI distinguish which differences actually matter?

Experts observe by choosing which differences matter (qualitative judgment); AI finds patterns and probabilities (quantitative). AI generates text from prompts without observing context, audience needs, or knowledge states—producing fabrication that mimics observation's form without its epistemic process.

Does AI separate intellectual form from the thinking behind it?

Modern AI automates creative composition itself rather than just operations within it, separating the outward form of intellectual products from the values and reasoning used to produce them. This mechanism allows exchange value to float free from use value.

Is the LLM a tool or a new form of intelligence itself?

Following McLuhan's logic, the model's cultural impact comes from its medium-properties—making intelligence generative and liquid—not from transmitting pre-existing intelligence. The model constitutes intelligence rather than delivering it.

Why does AI output change with every prompt and context?

AI outputs exhibit essential mutability—they vary with sampling, prompt wording, and audience interpretation. This is not a defect but a defining feature of tokens as media, making them fundamentally different from fixed commodities and resistant to traditional quality assurance.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.