Per Claude code review: removing the Triton mm_plus_mm template on XPU
broke test_autotune_gemm_choice_validation which asserts TritonTemplateCaller
is always present when max_autotune=True. Guard the assertion for the
mm_plus_mm+XPU case.
Also add TODO(#184490) in mm_plus_mm.py to mark the XPU guard as temporary
— it should be lifted once the underlying Triton codegen accuracy bug is
fixed so XPU can benefit from the fused template perf.
Co-authored-by: Copilot [email protected]
SOCIAL SHARE CARD GENERATOR