Skip to main content

Practice03

Quantum benchmarking.

Longitudinal noise-model validation on IBM Quantum hardware: a fixed battery of shallow entangling circuits, run for months on the same device, asking a question the literature has not answered.

The research question

How long does a noise model stay good?

Today's quantum computers are noisy, and simulating that noise well matters for everything built on them. The standard approach derives a noise model from the device's published calibration data. Ours asks: for shallow entangling circuits on a fixed superconducting device, is calibration-model prediction error dominated by model form or by parameter staleness? And can a learned noise model beat the calibration models — and retain that edge when predicting future hardware behaviour?

That forward-transfer angle — how long a noise model stays good, its half-life — is the unclaimed niche, verified against the existing literature before we committed to it. The instrument is deliberately modest: Bell, GHZ, W, and product states on four qubits, measured in five tomography bases, submitted on IBM's free tier. Rigour over scale.

Early results

Three runs, seven months, battery frozen.

Every future run stays comparable, because the circuits do not change.

01

Calibration models are ~3× pessimistic

Chip-averaged calibration noise models predict far more error than the hardware actually produces on shallow entangling circuits. The published spec sheet and the device you get are different machines.

02

Per-qubit models beat noiseless

Models built from per-qubit calibration data predict real GHZ-state output with a Hellinger distance of 0.043 — better than the 0.055 the vendor publishes for comparable hardware, and better than assuming no noise at all.

03

No staleness penalty at four months

The surprise so far: noise-model predictions did not measurably degrade across a four-month gap between runs. If that holds, "how often must you recalibrate?" has a cheaper answer than the field assumes.

Discipline

Run like an engagement, not a hobby.

01

Pre-registered protocol

Hypotheses, metrics, and the analysis plan were fixed in writing before the later data points — the protocol can be amended, never rewritten. The same rule we apply to client test plans: decide what passing looks like before you run the test.

02

Every claim traces to a receipt

All results trace back to a recorded job ID and archived calibration snapshot. Three retrospective hardware runs are analysed so far, with the battery frozen so future runs stay comparable.

03

An honest ceiling

This is not peer-reviewed work, and we say so. The stated ambition is an arXiv note or workshop paper after at least six monthly runs — no sooner, because the forward-transfer question needs the time series.

Next step

Research discipline, applied to your system.

The habits that make research honest — fix the success criteria first, trace every claim to a receipt, state the ceiling — are the same habits that make a payroll cutover boring.

Start a conversation

We reply within one business day