Google XProf Unlocks TPU Kernel Secrets

Alps Wang

Alps Wang

Sep 23, 2026 · 1 views

Diving Deep into TPU Performance

Google's addition of cycle-level kernel profiling to XProf represents a substantial leap forward for developers working with custom Pallas kernels on TPUs. Previously, these highly optimized kernels were treated as opaque blocks, making fine-grained performance tuning challenging. The ability to inspect hardware performance counters at a granular level, down to individual clock cycles and specific hardware units like the MXU, scalar ALUs, and vector units, provides unprecedented visibility. This is particularly crucial because custom compilation paths (Pallas, Mosaic, Triton) bypass standard XLA optimization passes, rendering traditional static cost models less reliable. The introduction of runtime telemetry, with its 1µs resolution and an external event-triggered mode for sub-microsecond capture, directly addresses this by offering ground truth metrics from hardware registers. This allows developers to move beyond potentially inaccurate XLA estimates and anchor their optimizations on empirical data, leading to tangible performance gains, as demonstrated by the matmul example reducing kernel time by 30% through triple buffering.

However, the practical application of this powerful profiling suite comes with its own set of considerations. The limited budget of 4x28 counters per core necessitates careful selection based on the specific investigation, meaning teams must strategically choose which metrics to monitor for each optimization effort. While the documentation focuses on TPU v7 (Ironwood), the coverage for earlier TPU generations is not guaranteed, potentially limiting its immediate applicability for all users. Furthermore, the article highlights a "hierarchy of trust" where hardware register reads are ground truth, implying that XLA-derived metrics for custom kernels should be treated with caution. This shift in optimization strategy requires developers to adapt their mindset and debugging workflows. The availability of a dedicated Perf Counters View with over 16,000 raw counters is a boon for deep dives, but managing and interpreting such a vast amount of data will require expertise and potentially new tooling or analysis techniques. The demonstration, while illustrative, is based on a single Google-authored demo kernel, underscoring the need for broader community adoption and testing to understand the full spectrum of potential gains across diverse workloads.

This enhancement is invaluable for AI/ML engineers and researchers focused on maximizing the performance of their models on Google's TPUs, especially those utilizing custom kernel languages like Pallas. It empowers them to identify and resolve performance bottlenecks that were previously hidden, leading to more efficient model training and inference. By providing a more accurate and granular view of hardware utilization, XProf's new profiling capabilities enable developers to push the boundaries of what's possible with AI hardware. The emphasis on grounding optimizations in hardware-level counters rather than relying solely on compiler estimates is a critical paradigm shift that will drive more effective performance tuning in the long run. This move also aligns with the broader trend in high-performance computing towards more detailed, hardware-aware profiling for achieving peak efficiency.

Key Points

  • Google has added cycle-level kernel profiling to XProf for TPU workloads.
  • Developers can now see detailed performance insights for custom Pallas kernels, which were previously opaque.
  • The new suite samples hardware performance counters with 1µs resolution and an external event-triggered mode for sub-microsecond capture.
  • This allows for more accurate performance attribution and optimization by grounding analysis in hardware register data.
  • A matmul example demonstrated a 30% reduction in kernel time by optimizing memory stalls through triple buffering.
  • The profiling works at three levels: compiler inspection, static execution analysis (LLO), and runtime telemetry.
  • Limitations include a counter budget of 4x28 per core and potential lack of coverage on older TPU generations.
  • Developers are advised to rely on hardware counters over XLA-derived estimates for custom kernels.

Article Image


📖 Source: Google Adds Cycle-Level Kernel Profiling to XProf

Related Articles

Comments (0)

No comments yet. Be the first to comment!