“Understanding human intelligence” sounds innocuous, but will eventually unlock an ability to run human-brain-like algorithms on computer chips. I’ll argue that such algorithms will be very different from today’s LLMs, amounting to a “new intelligent species”—a species which will eventually vastly outnumber humans, think much faster than us, and be far more insightful, resourceful, and experienced than us. We’d better make sure it’s a species that we want to share the planet with, and which wants to share the planet with us!
This challenge has many facets, but my focus will be “the technical alignment problem”: what code could people write, such that the resulting brain-like Artificial General Intelligence (AGI) would feel intrinsically motivated by human welfare, norms, instructions, etc., rather than feeling callous indifference? I’ll argue that this problem is hard, unsolved, and mostly orthogonal to the process of “understanding human intelligence”.
I’ll suggest that the brain has something akin to a reinforcement learning reward function, which says that pain is bad, eating-when-hungry is good, etc. This reward function is centered around the hypothalamus and brainstem, and I’ll argue that all human desires—even “higher” desires for things like compassion and justice—come directly or indirectly from that innate reward function. If future programmers build brain-like AGI, there will likewise be a reward function slot in the source code, in which the programmers can put whatever they want. If they put the wrong thing, they’ll wind up with “high-functioning sociopath AGI”, pursuing unintended goals with ruthless ingenuity. Are there good reward functions which avoid that problem, and if so, what do they look like? That’s an open technical problem, but I will review some ideas and research directions, including better understanding human social instincts.
(40-minute talk + Q&A, intended for a general audience)