Modern generative models have enabled robots to acquire challenging dexterous skills directly from human demonstrations. Yet, even the state-of-the-art learned policies, such as TRI’s Large Behavior Model (LBM), often suffer unexpected failures during deployment.