Protect What Pays: JIJ at ICCAD 2026

2026/10/2Technical Blog

Our work on shot-budget-aware error detection has been accepted to IEEE/ACM ICCAD 2026 and will be presented as an invited paper and talk in the special session "Design Automation for Early Fault-Tolerant Quantum Computing," November 8–12 in San Jose.

We are pleased to share that "Protect What Pays: Shot-Budget-Aware Placement of Error Detection in Quantum Optimization Circuits" will appear as an invited paper at ICCAD 2026, the IEEE/ACM International Conference on Computer-Aided Design, with an accompanying talk in the special session Design Automation for Early Fault-Tolerant Quantum Computing at the San Jose Marriott.

The session, co-organized by JIJ's Global R&D Manager, Louis Chen, together with Anupam Chattopadhyay (Nanyang Technological University), Zhiding Liang (The Chinese University of Hong Kong) and Shigeru Yamashita (Ritsumeikan University), starts from a simple observation: as quantum computing moves from the NISQ era into the early fault-tolerant regime, the field needs cross-layer design-automation methods for resource estimation, physical modeling, placement, routing, timing analysis and calibration. Our paper is a worked example of exactly that — it treats error detection not as a physics question but as a compiler decision, with a cost model, an exact scheduling algorithm and a decision procedure that a toolchain can run before a single shot is executed.

The paper is by Kuan-Cheng Chen and Marcelin Gallezot of JIJ, with Samuel Yen-Chi Chen of Wells Fargo.

The question

Quantum optimization algorithms such as QAOA are run thousands or millions of times, and on today's hardware a good fraction of those runs are corrupted by noise. Error detection codes offer a cheap remedy: encode the circuit, insert a few syndrome checks, and throw away any shot in which a check fires. The Iceberg code popular on trapped-ion machines does this with only two extra qubits, and recent experiments have shown it improving the fidelity of QAOA.

But fidelity is not what a user pays for. Every check costs entangling gates, mid-circuit measurements and idle time; every rejected shot costs the time already spent on it; the encoded mixer is serial where the unencoded one is parallel. What the user actually buys is wall-clock: the time to observe a good bitstring, or the time to estimate the cost function to a target precision, within a fixed budget of QPU time.

Our paper asks the question in those terms. When does protection pay, how many checks should there be, where should they go, and when is the honest answer "do not encode at all"?

What we built

The core of the paper is a cost model that is exact where it matters. Every logical operation of the Iceberg code commutes with both of its stabilizers, so the syndrome a fault produces depends only on the fault, not on where it happened or on any of QAOA's continuous rotation angles. That single observation turns the acceptance probability, the clean-shot yield and the expected shot duration with early abort into closed forms with one calibrated constant, where the comparable literature model fits three.

On top of it we prove a regime-separation theorem: the two things users actually do with a quantum optimizer have different optima, and they are not close. For sampling workloads — run until you see a good bitstring — the acceptance rate drops out of the cost entirely, so a check can only help by aborting doomed shots early. For estimation workloads — measure the cost function to a target precision — rejected shots are pure loss, and the penalty for throwing them away grows faster than the saving from a cleaner ensemble. On one LABS circuit the three natural objectives put their optima at 0, 8 and 13 checks respectively. Optimizing the wrong one is not a small error.

We call the resulting pass Shotwise. It tabulates the segment statistics once, places s checks by an exact dynamic program (scalar for sampling, a Pareto program for estimation), scores each schedule by the user's objective, and encodes only if the best schedule beats the unencoded circuit. Placement is front-loaded, which the theory predicts: a check earns its place only when the shots it kills were going to be expensive.

What we found

Three results stand out, all under Quantinuum H2-1 error rates and a calibrated time model (4.07 ms per entangling gate, a mid-circuit measurement worth 49 gate times, one week of QPU time).

  • A one-number design rule. Whether encoding pays is governed by γ, the number of cost-layer entangling gates per qubit, which you can read off the problem graph. Estimation breaks even at γ* ≈ 7 for 12 logical qubits, rising slowly with size. Three-regular MaxCut (γ = 1.5) never benefits; Sherrington–Kirkpatrick instances (γ = 5.5) sit on the margin; LABS, the problem on which QAOA's strongest advantage claims rest (γ = 25), benefits by 4.7–10.9× in our model.

  • Sampling never broke even. Across the entire sweep, encoding never reduced the time to observe a good bitstring: the two idle code qubits impose a yield penalty that early abort cannot recover. The right answer for every sparse instance we tested was do not encode.

  • Fidelity-optimal schedules are slow. Choosing the number of checks that maximizes fidelity, as current practice does, costs up to 93× in time-to-solution against the budget-aware schedule; uniform spacing of four checks costs up to 18.5×. Round count and the encode decision dominate; fine placement is second-order.

Validated at three levels, with CUDA-Q on NVIDIA GPU Cluster

A closed-form model is only as good as its validation, and a stabilizer simulator cannot see a non-Clifford circuit, so we checked every prediction at three levels of trust. The analytic level is the model itself. The proxy level is Stim, sampling the Clifford instance of every benchmark-size configuration: 288 configurations up to 20 logical qubits, agreeing with the model to 0.10% mean relative error against a 0.13% shot-noise floor. The full level is NVIDIA CUDA-Q, simulating the actual non-Clifford circuits at the optimized angles with the noise model attached gate for gate; exact density-matrix simulation covers 32 configurations up to 8 logical qubits, and the GPU trajectory backend on an NVIDIA H100 extends this to benchmark size, k = 12, where a density matrix would need 69 GB and days of CPU time.

The full level agrees with the model on every encode decision, and it taught us something the proxy could not: the standard white-noise assumption behind every estimation cost in this literature is conservative, with the accepted ensemble's effective fidelity exceeding the model's clean-shot weight by 0.6–9.8 points, because a fault is not white noise and a single X fault at readout leaves an effective fidelity of exactly 1 − 4/k rather than zero. This is the kind of result that only exists at GPU-enabled scale. Below about 14 physical qubits a CPU density matrix suffices; above about 40 no full simulation of a noisy non-Clifford circuit exists on any platform; in the 18–32 qubit window between them, the memory bandwidth of an H100 is what turns a week into hours, which makes encoded QAOA with mid-circuit measurement and early abort, we think, a good example of the logical-level workload CUDA-Q was built for.

Why it matters

The practical message is the one the session was convened around. Early fault-tolerant hardware will not make error mitigation free; it will make the trade-offs sharper, and those trade-offs belong in the compiler. Error detection costs time, the objective that matters is time-to-solution under a shot budget, and the decision can be made at compile time from the problem's gate density and the device's calibration.

For dense, high-value problems such as LABS, a budget-aware schedule turns protection into a real speedup for estimation workloads; for sparse problems the compiler should decline to encode, and say so.

This is exactly the layer JIJ builds. Our stack already carries a problem from JijModeling through OMMX as a stable intermediate representation into Qamomile, which compiles optimization problems into circuits and transpiles them across Qiskit, Amazon Braket and other backends — so a pass like Shotwise has a natural home: it reads the problem's gate density and the target device's calibration at compile time and decides, before any QPU time is spent, whether and how to encode. Combined with the resource-estimation and benchmarking work in Qamomile and JijZept Quantum, it moves error mitigation from a hand-tuned experimental choice into a budgeted, auditable decision in the toolchain.

Error detection is usually presented as a physics knob. We think it is a compiler decision — and once you measure it in the only currency a customer cares about, wall-clock time under a fixed shot budget, the right answer is often "do not encode," and the toolchain should be able to say so.

— Louis Chen, Global R&D Manager, JIJ

Join us at ICCAD

The talk is part of the special session Design Automation for Early Fault-Tolerant Quantum Computing at the San Jose Marriott. ICCAD 2026 runs November 8–12, 2026.

https://iccad.com/2026/design-automation-for-early-fault-tolerant-quantum-computing

If you are working on error-detected or error-corrected optimization and want to compare notes, come find us in San Jose.

CONTACT

Talk to us about operational challenges, adopting JijZept,or collaborating on quantum technology research.

Contact Us

↗

Making society computable
to unlock humanity's potential.

[JIJ Inc.]
 #403, CIC, 3-3-6 Shibaura, Minato-ku, Tokyo,108-0023, JAPAN


[JIJ Europe Ltd.]
 F25 Atlas Centre,
 Rutherford Appleton Laboratory,
 Harwell Campus OX11 0QX
 United Kingdom


[JIJ EU (JIJ GmbH) ]
 Beiersdorfstraße 12, 22529 Hamburg, Germany

Japan · United Kingdom · Germany · United States

© 2026 JIJ Inc.

Optimization · AI · Quantum computing