Per review from jianyizh and guangyey: removing the is_xpu fallback and
re-enabling the scatter_add decomposition causes atomic contention
performance regression with overlapping pooling windows. Revert the
decomposition removal and update the comment to document the performance
rationale (not precision — the kernel fix from torch-xpu-ops #3765 is
landed and pinned).
The fallback stays for performance. Tests reverted accordingly:
- test_torchinductor_codegen_dynamic_shapes: keep TestFailure entries
(no Triton kernel generated while fallback is active) - test_torchinductor_opinfo: keep inductor_one_sample skip
Co-authored-by: Copilot [email protected]
SOCIAL SHARE CARD GENERATOR