TOPIC
Foundation Models
A subject the collection covers, read through 3 synthesis notes.
View as
Do different AI models actually produce diverse outputs?
Explores whether using multiple different language models together creates genuine diversity or whether shared training and alignment cause them to converge on similar answers despite independence.
Does polished AI output trick audiences into trusting it? Why do LLMs generate novel ideas from narrow ranges? Why do preference models favor surface features over substance? Why do multi-agent LLM systems converge without genuine deliberation? Does high-frequency text homogenize user input before generation?
Polished AI artifacts exploit professional appearance to simulate expertise LLM research ideation collapses into narrow clusters despite high novelty Preference models systematically favor five surface features humans reject Multi-agent LLM systems fail through silent agreement in over 60 percent of iterations High-frequency text homogenizes input through iterative user rephrasing toward model preference
What happens when models train on AI-generated content recursively? Why do different AI models generate similar outputs independently? Why do different language models independently produce similar outputs? Can AI output be genuinely novel or only at the margins? Do AI-generated posts crowd out human voices without any coordination or intent? Why do multiple language models independently produce similar outputs in influence campaigns? Why does RLHF alignment reduce the diversity of viewpoints in AI output? What happens to solidarity and community signaling when AI smooths out voice differences? Can few-shot examples narrow generative diversity in creative tasks? Why do sigmoid conflict curves look the same across different language models? Does optimizing directly for semantic diversity improve both reasoning quality and exploration? Does alignment training create bidirectional instruction and response mappings?See all 93 inquiring lines on this note →
Can deep learning theory unify around training dynamics?
Is learning mechanics—focused on average-case predictions and training dynamics rather than worst-case bounds—the emerging framework that finally unifies fragmented deep learning theory?
Language models show three hierarchical tiers of understanding with distinct mechanisms Marr's three levels provide structured methodology for LLM interpretability Human-parseable theory is essential for AI safety oversight regardless of AI self-understanding Shannon and Kolmogorov measures fail to value data for bounded learners Epiplexity measures learnable structure independent of computational resources needed
What stability techniques prevent collapse in policy-critic adversarial training? How does mechanistic interpretability complement learning mechanics in explaining deep learning? Why should deep learning theory prioritize average-case over worst-case analysis? Which hyperparameter theories best explain universal behaviors across neural networks? What solvable idealized settings reveal fundamental phenomena in realistic deep learning? How do classical mechanics and statistical mechanics provide methodological templates for learning theory? How does the Learning Law explain why all examples should contribute equally? Why do optimal learning dynamics improve scaling law coefficients specifically? What distinguishes surface mechanisms from the training regimes that produce them? How do learning dynamics on one example shift predictions on other responses?
Can humans understand deep learning before AI does?
Explores whether investing in human-parseable deep learning theory remains valuable even if AI systems eventually develop their own self-understanding. Centers on why this matters for safety oversight.
Can deep learning theory unify around training dynamics? Can we monitor AI reasoning without destroying what makes it readable? Does incremental AI replacement erode human influence over society? What happens to social order when AI removes ritual constraints? Can AI research itself without losing human oversight?
Learning mechanics unifies deep learning theory through training dynamics Optimizing reasoning traces for safety produces obfuscation instead of alignment Incremental AI adoption degrades implicit alignment by replacing human labor dependency. AI conversation lacks the ritual machinery that sustains human social order. AI can automate insight distillation and prior injection in research loops
Why do human-designed neural architectures eventually get replaced by learned ones? What makes a neural network circuit actually interpretable to humans? What does a human-parseable framework for deep learning look like? Why does human-AI collaboration preserve safety compared to autonomous self-improvement?