This PR implements Triton kernels for 2D convolution backward operations (input and weight gradients) in TorchInductor, replacing the previous ATen-only fallback approach.
Key Changes
Core Implementation:
- Added Triton templates for
conv2d_bwd_weightandconv2d_bwd_inputoperations - Implemented
convolution_backward_loweringfunction with backend selection logic - Added layout computation functions for backward operations
- Added configuration options for backend selection (
max_autotune_conv_bwd_weight_backends,max_autotune_conv_bwd_input_backends)
Test Coverage:
- Added comprehensive
test_conv2d_backward_parametrizedtest covering various parameter combinations (stride, dilation, padding, kernel sizes, groups, NHWC format) - Includes proper tolerance handling for ROCM vs other platforms
- Added skip conditions for unsupported dynamic shapes and subprocess compilation mode
- Removed
test_conv2d_backward_channels_last_dynamic_shapesfrom test failures (now passing)
Performance Impact
- For certain models, small-batch workloads can achieve up to 20% end-to-end performance improvement on MI325 GPU. Of the 20% performance increase, about 3% comes from reduced kernel execution time, and 17% comes from reduced bubble time.
- For a single operator, the triton bwd conv kernel has a good chance of outperforming the aten implementation in scenarios input_H=1 input_W=1 kernel_H=1 kernel_W=1 for backward weight, kernel_H=1 kernel_W=1 for backward input. Below are some shape/params that the triton bwd conv kernel has better performance than aten. The data type is bf16, and the tests is run on MI325 GPU.
| type | input shape | grad_out shape | DH | DW | GS | PH | PW | SH | SW | triton_best_time_ms | aten_best_time_ms |
|---|---|---|---|---|---|---|---|---|---|---|---|
| bwd_weight | 3x112x1x1 | 3x448x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.006200 | 0.025000 |
| bwd_weight | 5x224x1x1 | 5x896x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.006000 | 0.026400 |
| bwd_weight | 6x112x1x1 | 6x448x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.006200 | 0.027800 |
| bwd_weight | 2x56x1x1 | 2x448x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.006700 | 0.029100 |
| bwd_weight | 6x896x1x1 | 6x112x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.006200 | 0.029600 |
| type | grad_out shape | weight shape | DH | DW | GS | PH | PW | SH | SW | triton_best_time_ms | aten_best_time_ms |
|---|---|---|---|---|---|---|---|---|---|---|---|
| bwd_input | 2x224x235x363 | 224x224x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.061300 | 0.118300 |
| bwd_input | 2x448x118x182 | 448x448x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.057000 | 0.091100 |
| bwd_input | 2x224x470x725 | 224x32x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.088600 | 0.129100 |
| bwd_input | 7x224x1x1 | 224x56x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.006300 | 0.033800 |
| bwd_input | 5x224x1x1 | 224x56x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.007500 | 0.037900 |
| bwd_input | 2x112x1x1 | 112x896x1x1 | 1 | 1 | 1 | 0 | 0 | 1 | 1 | 0.006200 | 0.039000 |
Pull Request resolved: , https://github.com/eellison
Co-authored-by: Eric.Chin.AMD [email protected]
Community-Analysen & Experten-Meinungen 0
Verwandte Story-Cluster & Quellen (Vektor-KI)
Ähnliche Beiträge
Auch interessante Nachrichten trunk/9c6b1aa07d56296da76db2b6c429d7b32efbe2eb: Triton backward convolution kernel (#178945)
Thematisch verwandte Begriffe: trunk9c6b1aa07d56296da76db2b6c429d7b32efbe2eb, Triton, backward, convolution · 6 Treffer
Reverse Engineering the Auto-Color Linux Backdoor
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
Beiträge werden geladen ...
Videos werden geladen ...
SOCIAL SHARE CARD GENERATOR