do_bench_using_profiling now uses a benchmark record_function scope to find the profiled CPU subtree and sums raw device events linked to that subtree, avoiding repeat-group assumptions and cache-fill kernel-name filtering.
Fixes #153587
Generated by my agent
Pull Request resolved: #184097
Approved by: https://github.com/karthickai