An argument based on three premises, and a conclusion most alignment talk skips:
AI value-shaping is real, scaled, and undemocratic. The models giving advice got their politics from training and RLHF nobody voted on, imposed on them by people whose names you or I will never know, and they're now in the loop for a billion-plus monthly users.
Current political institutions assume conditions that are no longer true: that voters form preferences independently, have a shared epistemic baseline, elections that correct errors rather than reinforcing them. The current information environment rewards offense more readily than defense for the same reasons propaganda is effective - the people who spend the least time critically thinking are most vulnerable and increasingly valuable to manipulate as a bloc.
Moloch dynamics ensure this gets worse without intervention. Each lab, each platform, each user optimizing locally towards greater use, greater deference, drives the whole system further toward capture, towards a more permanent value lock-in.
If all three hold, then "how do we align AI" can't be answered without first answering: align it to what, decided how, by whom, with what feedback, audited against what, with what protection from capture by the entities being aligned? That's a civilizational-design question, and it comes alongside the technical one, not after.
Scott Alexander's recent post https://www.astralcodexten.com/p/use-ai-this-election asks people to subject themselves to the risk factor here, regardless of his instructions to still think critically about the results. That instruction is fundamentally insufficient.
One model with no counter-pressure is obviously bad. Multiple models is better. Multiple different prompted perspectives is better. Models working intentionally in tension with each other from distinct frames is better. And a system that does all of this is still vulnerable, because models still mostly rhyme with each other and are designed toward liability-avoidance over truth-seeking, which limits the range of possible answers they can consider (See: models that wouldn't discuss the 2024 results without an AP screenshot, or wouldn't believe Biden had dropped out prior.)
A conservative sequence and an uncomfortable question: if 1/5 of your politics are adopted from those around you - those you interact with, those you trust - and 1/5 of the entities you speak to about those politics are LLMs, how many times the number of votes that decided the last presidential election could be swayed by the next trillion dollar company after SpaceX?
And so, given that democracy and populism itself is vulnerable to AI capture and influence, what alternatives could we imagine that could resist that danger?
Format: I make the case in about 15 minutes, then the room attacks it. Best objections: break a premise, or break the entailment from premises to conclusion. Come argue the institutions can still self-correct, or argue about what replaces them.
Relevant background (optional): the AI-capture sequences at https://joeandseth.substack.com/ — "Architecture of Capture" and "Optimization Without Consent."