Sqam

Released Q-Learning with Scalar Adjoint Matching (SQAM), a cheaper off-policy RL algorithm for flow policies. arXiv · blog · code