Recent models are able to leverage increasing amounts of inference compute and are able to operate effectively over ever longer horizons. In this talk I will argue that, because of this, the field should fundamentally change how models are evaluated and especially how their dangerous capabilities are measured.