Tools
qcr:2609.13793.2

Backline

Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time quantum error correction (QEC) demand performance that traditional interpreted environments cannot provide. FPGAs and ASICs play a central role at these layers, but their specialized programming models make development rigid and time-consuming. CPUs, GPUs, and other accelerators introduce a different challenge: as infrastructure becomes increasingly heterogeneous, programming across different devices and their associated abstractions becomes more complex. Allowing researchers to write workloads in high-level languages that map to low-latency execution across diverse distributed target platforms will enable the development of key infrastructure for utility-scale quantum systems. For this, we introduce Backline, a heterogeneous compilation and runtime framework built within PennyLane and Catalyst. Backline allows us to design and build quantum-classical workloads for high-performance and low-latency devices, with compilation directly from a Python interface through MLIR. We demonstrate the compilation and execution of several quantum workloads with low-latency data movement across a mix of CPUs, GPUs, and FPGAs, for both local and distributed remote hardware targets, all from a vendor-agnostic Python frontend. With an AMD VPK120 FPGA board as the controller, issuing each round from its hardware-handshake engine, we measured median steady-state round-trip latencies over RoCE v2 of 2.305 μs to an AMD Ryzen Threadripper PRO CPU and 4.5 μs to an AMD Instinct MI210 GPU across 106−1 rounds per path, demonstrating microsecond-scale synchronous co-processing.
Compilation
Uploaded 2 weeks ago
23
Views
Citing this entry? Use this QCR ID
Uploaded by
AS
Antal Száva

Overview

PennyLaneAI/backline
151
README.md

Backline

Backline is an open platform for compilation and low-latency execution by Xanadu and AMD that dynamically connects quantum workloads to the right classical engine.

With PennyLane and Backline, anyone can write a QEC encoder or decoder from Python, test it with meaningful quantum algorithms, and immediately deploy it for near-real-time execution on CPUs, GPUs, FPGAs, and QPUs — while supporting the need to drop through abstractions and write increasingly optimized and low-level code.

[!NOTE] The core Backline implementation and source code lives natively within the PennyLane and Catalyst repositories. This repository holds the demonstrations, benchmarks, and the cross-build system accompanying the manuscript "Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs".

[!NOTE] Backline is currently under heavy development — if you have suggestions on the API or use-cases you'd like covered, please open a GitHub issue in the relevant repository (PennyLane or Catalyst), or reach out to [email protected] and [email protected]. We'd love to hear about how you're using the library, collaborate on development, or integrate additional devices and frontends.

Key Features

  • Single-digit microsecond latency: Achieve under 3-μs end-to-end loops. Backline treats CPUs and GPUs as highly responsive endpoints to support the tight co-processing needed for QEC backup decoding.

  • Scale from R&D to production: Prototype immediately on standard CPUs—bypassing the need for any GPUs. Then, seamlessly scale to consumer- and enterprise-grade GPUs and FPGAs from the exact same PennyLane application.

  • Python-native: Build entirely in Python. Write optimized GPU kernels using Triton and Gluon, or integrate pre-compiled libraries alongside your quantum logic. Easily jump through abstraction layers without switching frameworks.

  • Infrastructure agnostic by design: Leveraging the LLVM ecosystem, Backline supports CPUs, GPUs, FPGAs, and custom devices to meet the diverse error-correction needs of any quantum platform.

Getting started

Once Backline is installed, you can get started by checking out the Backline tutorial, then working your way through the demos in this repository. To reproduce the paper, run those and the benchmarks.

Also make sure to check out the technical documentation, technical manuscript, and Backline whitepaper.

Architectural overview

Backline provides the following three main abstractions for use with PennyLane and Catalyst:

  • Controllers: This is the classical hardware node (such as a CPU or FPGA) that controls the QPU (a quantum hardware or simulator qp.device), receives quantum measurement results, and initiates data transfers with other hardware devices (coprocessors). For example, it might perform QEC syndrome measurements on the QPU, and send these to a coprocessor for decoding.

  • Coprocessors: These are hardware device nodes (such as CPUs, GPUs, or FPGAs) that receive information from a controller for processing. They run specific coprocessing functions, potentially as a persistent kernel, such as a QEC decoder.

  • Backline: A representation of the complete hardware infrastructure supporting the quantum-classical program. The backline includes a controller, one or more coprocessors, and a transport method. A backline object is given directly to a QNode in place of a traditional QNode qp.device, and orchestrates the remote executor and the RDMA network the controllers and coprocessors talk over, separate from the network used to log into remote machines.

If you are an AI agent, read AGENTS.md first: the same material, plus the specific traps that have caught agents here before.

Repository Overview

This repository contains the benchmark data from the manuscript "Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs", as well as the cross-build system for reproducing the stack demonstrated.

In addition, a variety of demos are provided, highlighting the compilation, deployment, and execution of a single quantum error-corrected PennyLane program onto several machines at once with low-latency execution.

  • benchmarks: the latency measurements the manuscript reports.
  • config: the cross-build system (xbuild) and the machines to run on (machines.toml) to reproduce the results from the manuscript.
  • data: the measurements themselves, and the notebook to plot.
  • demos: demo programs, spanning from a single CPU machine to FPGA-to-GPU.
  • scripts: helpers for running the above.

Installation

Backline requires a recent version of PennyLane, Catalyst, and Lightning. We recommend installing version v0.46.0b1 for PennyLane and Lightning, v0.16.0b1 for Catalyst (either from source or using pre-built wheels for local demos).

To install Backline, please see INSTALL.md for instructions and requirements. Note that due to the wide range of system, network, and hardware configurations you can use Backline with, there are different installation requirements and steps depending on your needs:

Tier Features and usage Requirements Corresponding Demo
1 Local CPU-CPU interactions on any Linux CPU machine PennyLane and Catalyst 1
2 Local CPU-CPU interactions over RDMA As above, plus a device that supports the libibverbs interface; soft-RoCE will do, an RDMA NIC is optional 1a
3 Remote CPU-GPU interactions As above, plus a server containing a GPU and an RDMA NIC, accessible over SSH 2, 2a, 3
4 Remote CPU-FPGA or GPU-FPGA interactions As above, plus a Xilinx VPK120 board connected to the server via RDMA 4, 5, benchmarks

Once Backline is installed, you can verify your installation locally by compiling and executing a simple local CPU-CPU interactions on a Linux CPU machine via memcpy.

Two environment variables to be aware of when using Backline are:

  • CATALYST_ROOT: the local Catalyst build tree, default ~/catalyst.
  • BACKLINE_BUNDLES: the cross-built stacks the remote machines deploy, one directory per bundle name. Defaults to what config/xbuild publishes.

Authors

Backline is the work of many contributors.

If you are doing research using Backline and PennyLane, please cite our papers:

@misc{lee2026,
  title={Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs},
  author={Joseph K. L. Lee and Mehrdad Malekmohammadi and Hong-Sheng Zheng and Shuli Shu and Cheick Doumbia and Kalman Szenes and Mehran Zamani Abnili and Thomas Ainsworth and Matthew Seymour and Thomas Germain and Leonhard Neuhaus and Josh Izaac and Lee J. O'Riordan},
  year={2026},
  eprint={2609.09270},
  archivePrefix={arXiv},
  primaryClass={quant-ph},
  url={https://arxiv.org/abs/2609.09270},
}

License and acknowledgements

Backline is free and open source, released under the Apache License, Version 2.0.

AMD, AMD Ryzen, AMD Ryzen Threadripper, AMD Instinct, AMD ROCm, Radeon, Versal, and Xilinx are trademarks of Advanced Micro Devices, Inc.

Join the Discussion

Comments (0)

No comments yet. Be the first to share your thoughts!

Publication

doi:10.48550/arxiv.2609.09270
Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs

Joseph K. L. Lee, Mehrdad Malekmohammadi, Hong-Sheng Zheng, Shuli Shu, Cheick Doumbia, Kalman Szenes, Mehran Zamani Abnili, Thomas Ainsworth, Matthew Seymour, Thomas Germain, Leonhard Neuhaus, Josh Izaac, Lee J. O'Riordan

Versions

v2 Latest
Sep 10, 2026
qcr:2609.13793.2

Cite all versions? Use the base QCR ID to always reference the latest version of this entry.

Keywords

gpu
qec
heterogeneous compilation
mlir
fpga

You may also like5