trunk/40df7254415ed93bfe83b9d406b44f171ecc5479: Fix flex_attention score_mod with no score gradient (#185991)
🔒
https://github.com
«FlexAttention builds a joint graph for score_mod so the backward template can compute the gradient of the modified scores with respect to the raw attention scores. When score_mod returns a value that is independent of sc...»
Automatische Weiterleitung...
1.5s