Posted by Ashraf Bhuiyan, AG Ramesh from Intel, Penporn Koanantakool from Google
and 8-bit low precision data types. Low precision data types are widely used and provide significant improvement over the default 32-bit floating format without significant loss in accuracy.
We are happy to announce that these features are now available as a preview in the nightly build of TensorFlow on published by Intel.
How to take advantage of AMX optimizations on 4th Gen Intel Xeon.
Intel AMX optimizations are included in the official TensorFlow includes preliminary support, however full support will be available in a subsequent stable release.
Users running TensorFlow on Intel 4th gen Intel Xeon can take advantage of the optimizations with minimal changes:
a) For bfloat16 mixed precision, developers can accelerate their models using Keras mixed precision API, as explained , for example RN50, BERT, SSD-RN34 that have been previously quantized with and the launch_benchmark script from![]() |
Here in the chart, inference with mixed precision models on a 4th Gen Intel Xeon was 1.9x to 9.6x faster than FP32 models on a 3rd Gen Intel Xeon. (BS=x indicates a large batch size, depending on the model)
![]() |
Training models with auto-mixed-precision on a 4th Gen Intel Xeon was 2.3x to 5.5x faster than FP32 models on a 3rd Gen Intel Xeon.
![]() |
Similarly, quantized model inference on a 4th Gen Intel Xeon was 3.3x to 19x faster than FP32 precision on a 3rd Gen Intel Xeon.
In addition to the above popular models, we have tested 100s of other models to ensure that the performance gain is observed across the board.
Next Steps
We are working to continuously tune and improve the Intel AMX optimizations in future releases of TensorFlow. We encourage users to optimize their AI models with Intel AMX on Intel 4th Gen processors to get a significant performance boost; not just for inference, but also for pre-training, fine tuning and transfer learning. We would like to hear from you, please provide feedback through the .
Acknowledgements
The results presented in this blog is the work of many people including the TensorFlow and oneDNN teams at Intel and our collaborators in Google’s TensorFlow team.
From Intel: Md Faijul Amin, Mahmoud Abuzaina, Gauri Deshpande, Ashiq Imran, Kanvi Khanna, Geetanjali Krishna, Sachin Muradi, Srinivasan Narayanamoorthy, Bhavani Subramanian, Yimei Sun, Om Thakkar, Jojimon Varghese, Tatyana Primak, Shamima Najnin, Mona Minakshi, Haihao Shen, Shufan Wu, Feng Tian, Chandan Damannagari.
From Google: Eugene Zhulenev, Antonio Sanchez, Emilio Cota.



SOCIAL SHARE CARD GENERATOR