# Q-Day Index — methodology snapshot

Generated: 2026-08-27

This file is the concatenation of the project's reviewed methodology documents, in full and unedited. It is the citable form of the methodology page.



---

<!-- event-definition.md -->

# Event Definition

Version: 1.1 · Status: Part XII items 1-4 approved by founder (2026-07-21), unchanged from
this draft — see [decision-log.md](decision-log.md) · Last updated: 2026-07-21

See also: [product-definition.md](product-definition.md), [model-methodology.md](model-methodology.md).
Source of truth: `IMPLEMENTATION_PLAN.md` Part I, section 1.

## Terminology

Use the established term **Cryptanalytically Relevant Quantum Computer (CRQC)**: a quantum
system capable of using quantum algorithms to break a cryptosystem that is secure against
classical computers. The public-facing name stays **Q-Day**, but the methodology defines it in
relation to a CRQC — the public name is marketing shorthand, not a separate technical concept.

## Primary technical event — ECC-256 capability

> The earliest date on which a fault-tolerant quantum system can recover a private key from a
> specified, practically deployed 256-bit elliptic-curve target within a declared attack-runtime
> scenario and success probability.

Initial targets: NIST P-256, secp256k1, and optionally a separate ~128-bit-classical-security
curve target. These are modelled independently — they are not combined until the resource model
proves the comparison is valid.

## Secondary technical event — RSA-2048 capability

> The earliest date on which a fault-tolerant quantum system can factor a specified RSA-2048
> modulus within a declared attack-runtime scenario and success probability.

## No "general-purpose" requirement

The security event is the existence of *any* credible quantum system capable of completing the
defined attack. A general-purpose requirement could incorrectly exclude a specialised
cryptanalytic architecture built for exactly this attack.

## Runtime scenarios

Distinct, separately modelled runtime scenarios rather than one vague "operationally
meaningful" window:

| Scenario | Window |
|---|---|
| `rapid` | 1 hour |
| `transaction_window` | 10 minutes (where relevant, e.g. crypto mempool exposure) |
| `operational` | 24 hours |
| `extended` | 7 days |
| `strategic` | 30 days |

Public launch selects one primary scenario, disclosed next to the P10–P90 hero; the others
remain visible in methodology and sensitivity analysis. **Part XII item 4, resolved
2026-07-21: `operational` (24 hours) is the primary scenario.**

## Three forecast outputs

1. `public_evidence_capability_date` — forecast from public, validated technical evidence only.
2. `hidden_adjusted_capability_range` — sensitivity range showing how undisclosed-capability
   assumptions alter the baseline. See [hidden-capability-model.md](hidden-capability-model.md).
3. `public_confirmation_date` — optional disclosure scenario; must not be presented as more
   objective than the capability date.

## Primary hero

Full public-evidence forecast interval, always:

```text
P10 — P50 — P90
```

Hero must show: P10, P50 (marker inside the interval, not dominant), P90, the controlling
cryptographic target, the selected runtime scenario, model version, last publication date, and
movement since the previous version. Hidden-capability results appear as a separate sensitivity
band — never silently blended into the public-evidence baseline.

An optional compact secondary countdown (e.g. `7 years · 4 months · 18 days`, no seconds,
visually subordinate, labelled as derived from the weekly model version) may sit below the
uncertainty hero. No continuously ticking seconds display, ever.

## Required public disclaimer

> This is a probabilistic model of CRQC capability based on public evidence. Its primary result
> is a forecast interval built from the percentiles the evidence can resolve. A percentile that
> does not cross within the model's horizon is reported as undefined, never estimated, together
> with the share of scenarios that do not cross. The model is recalculated weekly and does not
> claim day-level or second-level certainty.

If a compact countdown is shown, add:

> The compact countdown is derived from the published P50 date and is a secondary visual aid,
> not the model's primary scientific output. A compact countdown is shown only when a median
> resolves.

### Amendment, 2026-07-27 (founder decision, issue #84)

**What changed.** The disclaimer previously read *"The primary result is the P10–P90 forecast range,
with P50 shown as the median estimate"*, and the countdown clause described the countdown as though
a median always existed.

**Why.** Censoring runs top-down, so a run can resolve P10 while P50 and P90 do not. The wording
above was resolved on 2026-07-21, when the model produced no crossings at all and that state could
not occur; it did not contemplate a state that did not yet exist. From the error-budget correction
(#72/#73) onward it does occur — the live run resolves P10 only — so the required text asserted a
range and a median on the same screen as a hero explaining why there is none. The amendment states
the model's **output contract** rather than one of its outcomes, which is what made the old wording
wrong the moment the outcome changed.

**Which direction it moved.** The disclaimer became wrong because the model got *better* at
reporting what it cannot resolve, not worse. `reported as undefined, never estimated` is now part of
the required text, so that refusal is a contract rather than a presentational choice.

**Why this note exists.** `public.views.methodology_snapshot` serves the live markdown from
`METHODOLOGY_DOCS_ROOT` rather than a copy captured at publication, so a citation of an older model
version resolves to whatever this file says today. Without a dated note the amendment would quietly
rewrite what past versions appear to have said — the opposite of what the model-history page exists
to do. Whether the §52 snapshot should be captured per published version instead of served live is a
separate and larger question, filed on its own.

## Earliest-target rule

```text
min(ECC target date, RSA-2048 target date)
```

is valid as the overall published estimate only if each target is independently modelled,
runtime scenarios are comparable, uncertainty is preserved, and the UI states which target
currently determines the published forecast. **Part XII item 2, resolved 2026-07-21: the
public forecast targets ECC-256 only** — this earliest-target rule is not currently exercised
(no RSA blending), kept here as documentation of the rule that would apply if that ever changes.

## Resolved questions (Part XII items 1-4, 2026-07-21)

All four approved unchanged from this draft — see [decision-log.md](decision-log.md) for the
full record:

- Item 1: Q-Day defined as the CRQC-based technical event above.
- Item 2: ECC-only (not earliest-of-ECC-and-RSA) is the published headline.
- Item 3: ECC target configuration is NIST P-256 + secp256k1.
- Item 4: `operational` (24 hours) is the primary runtime scenario for the public hero.



---

<!-- model-methodology.md -->

# Forecast Model Methodology

Version: 1.0 (draft) · Status: draft, pending founder review (Part XII items 5-11) ·
Last updated: 2026-07-18

See also: [event-definition.md](event-definition.md), [evidence-taxonomy.md](evidence-taxonomy.md),
[hidden-capability-model.md](hidden-capability-model.md),
[weekly-publication-process.md](weekly-publication-process.md).
Source of truth: `IMPLEMENTATION_PLAN.md` Part IV.

## Architecture — four layers

### Layer A — architecture-specific structural models

Separate pathways, never aggregated into one raw "qubit growth" curve: superconducting, trapped
ion, neutral atom, photonic, topological/emerging, modular/networked. Each pathway models:
demonstrated physical scale, logical encoding/code family, logical qubit count, logical
error per operation/cycle, error-suppression scaling, syndrome-extraction cycle time,
fault-tolerant gate set, magic-state (or equivalent) resource production, decoder
latency/throughput, routing/connectivity overhead, uptime/calibration requirements,
manufacturing/control-system constraints. A logical qubit from one experiment is never treated
as equivalent to a logical qubit from another code/architecture without normalisation.

### Layer B — target-specific resource models

Separate resource-estimate records per target (P-256, secp256k1, RSA-2048) and per runtime
scenario, each recording: algorithm/paper version, logical qubits, gate counts by type, circuit
depth, assumed physical error rate, assumed QEC code, factory overhead, routing overhead,
runtime assumptions, success probability, physical-qubit estimate, uncertainty and omitted
engineering factors. Never combine "best" assumptions from different papers unless technically
compatible.

### Layer C — scenario and Monte Carlo engine

Per model run: select an immutable approved evidence snapshot → select one architecture pathway
per simulated competitor → sample correlated future trajectories → sample model-parameter
uncertainty → sample roadmap delivery reliability → sample algorithmic-resource improvements →
sample engineering delays → evaluate threshold crossing per target/runtime scenario → retain
architecture/target responsible for each crossing → calculate P10/P25/P50/P75/P90 → calculate
yearly cumulative probabilities → calculate sensitivity/uncertainty decomposition → save random
seeds and simulation artefacts. Correlated variables are never sampled independently.

### Layer D — expert priors and later Bayesian updating

v1: explicit, versioned distributions derived from public evidence, historical roadmaps, and
documented expert elicitation, called a "scenario-based probabilistic forecast" — not a claimed
complete Bayesian model unless likelihoods/posterior updates are formally implemented. Later
versions: define priors and likelihoods, update architecture-specific latent states, validate
posterior predictive behaviour, publish the update rules.

## Bottleneck and competing-risk logic

A valid attack requires ALL of: `logical_capacity AND logical_error_budget AND
non_clifford_resource_supply AND operation_depth AND decoder_throughput AND runtime AND
engineering_availability`. The overall Q-Day distribution is a competing-risk result across
hardware architectures, companies/labs, cryptographic targets, and runtime scenarios — public
diagnostic indices are never averaged to produce the forecast. Funding, patents, hiring, and
roadmaps stay contextual indicators until a documented calibration method exists.

## Weekly publication cycle

See [weekly-publication-process.md](weekly-publication-process.md) for the full cadence and
publication rule.

## Model version record

Every model run stores: `model_version, configuration_version, evidence_snapshot_id,
code_commit_sha, container_image_digest, random_seed, scenario_count, run_started_at,
run_completed_at, status, ecc_capability_p10/p50/p90, ecc_public_p10/p50/p90,
rsa_capability_p10/p50/p90, rsa_public_p10/p50/p90, probability_by_year, previous_version,
movement_days, primary_drivers, secondary_drivers, validation_report, published_at`, and — since
Stage 1 — `input_bundle`, the frozen input set the run consumed.

## Reproducibility, and a correction to what this document used to claim

**The claim now.** Every public forecast is reproducible from: code commit, model configuration,
random seed, container image, and the run's **`ForecastInputBundle`** — the record of every
numerical and configuration input the forecast engine actually consumed.

**The distinction that matters, because these are not two names for one thing.** *The evidence
snapshot freezes approved evidence and the source registry. The `ForecastInputBundle` freezes every
numerical and configuration input actually consumed by the forecast engine.* A run records both.
Only the second one lets the run be re-executed.

**Correction, 2026-07-30.** From Phase 11 until 2026-07-29 this section read:

> Every public forecast must be reproducible from: code commit, model configuration, evidence
> snapshot, random seed, container image (§17).

**That was not true, and a reader who cited an earlier version of this document should know it.**
Two independent reasons, both now fixed:

1. **The evidence snapshot did not feed the model.** `ModelInputSnapshotItem` has a foreign key to
   `Evidence` and to nothing else, while the forecast core read `HardwareCapability` and
   `ResourceEstimate` live from the database. A run's snapshot checksum could be perfectly valid
   while proving nothing about the figures that entered the forecast: editing a hardware row after
   a run changed what the same seed produced, and the recorded provenance did not move.
2. **One input was never controlled by anyone.** Competitor order came from a `SELECT DISTINCT`
   with no `ORDER BY` — the physical row order of the table — and each competitor draws its
   parameters in loop order. Two executions over identical rows were measured producing different
   orders (#138/#139). Re-running from a perfect record of everything else would not have
   recovered it.

Measured cost of the first defect: over one integration fixture, three ordinary row edits moved the
headline ECC-public median from 2025-08-31 to 2043-12-05 — **eighteen years** — between a run over
its frozen bundle and the same run over freshly-read rows. The full accounting, including Stage 1's
ten acceptance criteria and what each is satisfied by, is in
[decision-log.md](decision-log.md) under "Stage 1 complete".

**Runs that predate bundles are not reproducible and are not presented as such.** They carry
`legacy_reproducibility_status = incomplete_legacy`, their publication convergence check reports
`untested` with a stated reason rather than approximating it against today's rows, and Stage 5 will
block them from new headline publication. That statement is maintained in one place — see
[decision-log.md](decision-log.md), "Legacy runs".

## Forecast evaluation

Since Q-Day has not occurred, direct calibration of the overall event date is impossible.
Evaluate instead through: hindcasts using historical evidence cut-off dates; accuracy of
component forecasts (roadmap delivery, logical-error improvements); comparison with independent
expert surveys; sensitivity stability; forecast-revision behaviour; external scientific review;
publication of failed or materially changed assumptions. A passing test suite validates the
software, not the scientific forecast.

## Approved baselines — Part XII items 5, 6, 7, 11 (resolved 2026-07-22)

Founder-approved after a two-pass research and verification effort (candidate values sourced,
then cross-checked against primary texts before sign-off — several first-pass figures were
corrected or dropped in the process, noted inline below). These are the v1 baseline inputs;
each is versioned like any other model configuration and may be revised as newer papers or
hardware results appear — revision is a new configuration version, not a silent edit.

### Item 5 — Resource-estimate baselines

**RSA-2048 (factoring, Layer B):** Gidney, "How to factor 2048 bit RSA integers with less than a
million noisy qubits," arXiv:2505.15917 (2025) — 1,409 logical qubits (peak), &lt;1,000,000
physical qubits, 6.5×10⁹ Toffoli gates, &lt;1 week runtime, 0.1% uniform gate error rate, 1μs
surface-code cycle time. Supersedes the original canonical baseline, Gidney &amp; Ekerå, "How to
factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits," *Quantum* 5:433 (2021)/
arXiv:1905.09749 (~20,000,000 physical qubits, ~2.6×10⁹ Toffolis, 8 hours) — the 2025 paper
trades more Toffoli gates (and longer runtime) for a 20x cut in physical qubits, and is itself
~300x more gate-efficient than the intervening Chevignard, Fouque &amp; Schrottenloher, "Reducing
the Number of Qubits in Quantum Factoring," IACR eprint 2024/222 (CRYPTO 2024, ~2×10¹² Toffolis)
approach it builds on. Both papers are recorded as model versions; only the 2025 figures are the
active v1 baseline.

**ECC discrete log, generic 256-bit curve (Layer B):** Roetteler, Naehrig, Svore &amp; Lauter,
"Quantum Resource Estimates for Computing Elliptic Curve Discrete Logarithms," ASIACRYPT 2017/
arXiv:1706.06752 — **2,330 logical qubits** (n=256), 1.26×10¹¹ Toffoli gates, 1.16×10¹¹ Toffoli
depth; canonical logical-qubit-count baseline. **Corrected 2026-07-27** from "≈2,126", a figure that
appears in no primary text — see "Corrections" below. Also added as a separate, non-default record:
Häner, Jaques, Naehrig, Roetteler &amp; Soeken, "Improved quantum circuits for elliptic curve discrete
logarithms," arXiv:2001.09580 (2020) — 2,124 logical qubits, a T-depth-optimised circuit, kept
independently citable rather than merged.
Litinski, "How to compute a 256-bit elliptic curve private key with only 50 million Toffoli
gates," arXiv:2306.08585 (2023) — ~50 million Toffolis, the gate-optimized alternative
configuration (same target, different resource trade-off — never combined with the Roetteler
qubit count as a single "best of both" estimate, per this doc's existing rule).

**secp256k1-specific (Layer B, distinct from generic P-256):** Babbush, Zalcman, Gidney,
Broughton, Khattar, Neven (Google Quantum AI), Bergamaschi (UC Berkeley), Drake (Ethereum
Foundation), Boneh (Stanford), "Securing Elliptic Curve Cryptocurrencies against Quantum
Vulnerabilities: Resource Estimates and Mitigations," arXiv:2603.28846 (2026) — first
secp256k1-specific (not generic-256-bit-curve) estimate: low-qubit variant &lt;1,200 logical
qubits/&lt;90M Toffolis; low-gate variant &lt;1,450 logical qubits/&lt;70M Toffolis; &lt;500,000
physical qubits, ~9-23 minute runtime, superconducting surface code, 10⁻³ physical error rate.
Worth noting for governance: this result was disclosed via a zero-knowledge proof (SP1 zkVM +
Groth16 SNARK) rather than publishing the attack circuits directly — a responsible-disclosure
choice that affects independent verifiability and should be described as such in any public
methodology page, not presented as identically verifiable to an openly-published circuit.

Two candidate generic-256-bit-curve figures (~1,098 and ~1,333 logical qubits, seen in secondary
aggregator coverage during research) could not be traced to a resolvable primary citation after
a dedicated verification pass and are explicitly **not** included as baseline inputs — noted
here so a future session doesn't re-introduce them without first finding the actual paper.

### Item 6 — Hardware capability baselines (as of mid-2026)

| Org | Result | Physical qubits | Logical qubits | Source type |
|---|---|---|---|---|
| Google Quantum AI | "Quantum error correction below the surface code threshold," *Nature* (Dec 2024)/arXiv:2408.13687 — distance-7 surface code, 0.143%±0.003% logical error/cycle, Λ=2.14±0.02 suppression per distance+2 | 101 | 1 (memory qubit) | Peer-reviewed (Nature) |
| Quantinuum | Dasu, DeCross, Guo, Lavasani, Behrends, Benhemou, Chen, Mayer, Self, Simsek, Srivastava, et al. (46 authors), "Computing with many encoded logical qubits beyond break-even," arXiv:2602.22211 (Feb 2026) | 98 (Helios) | 94 (iceberg QED codes) / 48 (two-level concatenated QEC codes) | Preprint, Quantinuum |
| Atom Computing + Microsoft | Nov 2024 neutral-atom (ytterbium) demonstration | 1,180 | 24 (entangled) / up to 28 (with error detection/correction) | Company blog (Microsoft Azure Quantum) |
| IBM | **Confirmed gap, not an oversight**: flagship result is Bravyi, Cross, Gambetta, Maslov, Rall &amp; Yoder, "High-threshold and low-overhead fault-tolerant quantum memory," *Nature* 627:778-782 (2024) — a theoretical/architectural qLDPC code-design result (0.7% threshold, 1/24 encoding rate), **not executed on hardware**. Smaller IBM hardware papers exist but are single-logical-qubit, distance-3-to-9 repetition-code demonstrations only. Public roadmap targets: "Starling" (2029, 200 logical qubits/100M gates), "Blue Jay" (2,000 logical qubits/1B gates) | — | none demonstrated | Roadmap/theoretical only |

Cross-check worth recording: the ~10⁻³ physical error rates assumed in the Item 5 resource
estimates are in line with or slightly better than today's best demonstrated two-qubit gate
fidelities above — the gap between current hardware and a cryptographically-relevant machine is
in scaling qubit count and engineering, not per-gate physics we haven't demonstrated yet.

### Item 7 — Probability distributions

Global Risk Institute Quantum Threat Timeline Report is the sole recurring, rigorous
expert-elicitation benchmark in this space — no independent second survey of comparable rigor
exists (NIST/ETSI publish migration/standards guidance, not a comparable probability survey);
this absence of a cross-check is itself recorded as a model limitation, not silently omitted.

- **2025 edition** (Mosca &amp; Piani, evolutionQ Inc., published via Global Risk Institute, March
  2026), 26 experts surveyed (declining panel: 37 in 2023 → 32 in 2024 → 26 in 2025, itself a
  possible response-rate/selection caveat worth disclosing publicly): 10-year horizon (~2036)
  28%-49%; 15-year horizon (~2041) 51%-70% ("likely"); 20-year horizon (~2046) 92% of respondents
  at ≥50%.
- **2024 edition** (evolutionQ, published December 2024), 32 experts surveyed: 10-year horizon
  19%-34%.

Both editions are recorded as versioned model inputs (2024 and 2025 are distinct configuration
versions, never blended into one number). The public methodology pages must disclose the
declining-panel-size caveat and the absence of an independent second survey alongside the
figures themselves.

### Item 11 — Hindcasting methodology and minimum evidence archive

No quantum-computing-specific hindcasting precedent exists to adopt directly — Sevilla &amp;
Riedel, "Forecasting timelines of quantum computing," arXiv:2009.05045 (2020) is a forecasting
paper (not itself a rolling-cutoff self-validation study) and is the closest quantum-specific
analog found. The approved approach instead **adapts** a peer-reviewed, general-purpose
technology-forecasting design rather than presenting an invented methodology as an adopted
standard: Farmer &amp; Lafond, "How predictable is technological progress?," *Research Policy*
45(3):647-665 (2016)/arXiv:1502.05274 — a rolling-origin hindcast design (pick a historical
date, forecast forward using only data available as of that date, compare against what actually
happened, repeat across multiple technologies and horizons), originally validated across 53
technologies. This is recorded explicitly as a **novel methodology decision built on a borrowed
design pattern**, not an adopted external quantum-computing standard, and must be described that
way in any public methodology page. The minimum evidence archive needed before public launch
(how far back historical data must reach for a meaningful rolling-origin test) is an
implementation detail of adapting this design, not yet separately specified.

### Corrections — 2026-07-27 (founder-approved, two-pass verified against arXiv LaTeX sources)

**The logical-error budget was computed against circuit depth. No paper in this literature does
that.** Both baselines that carry a physical layer budget error against **spacetime volume** —
logical qubits multiplied by error-corrected rounds — and say so explicitly:

> Gidney 2025 (arXiv:2505.15917), §3.2: "each shot will take roughly 12 hours and involve fewer than
> 1600 logical qubits (including idle hot patches). Given the assumed surface code cycle time of 1
> microsecond, this implies 1600 · 12 · 60 · 60 · 10⁶ ≈ 6.9 · 10¹³ logical qubit rounds of runtime to
> protect. Choosing a target logical error rate of 10⁻¹⁵ **per logical qubit round** will thus result
> in a no-logical-error shot rate of (1 − 10⁻¹⁵)^(6.9·10¹³) ≈ 93.3%."

> Litinski 2023 (arXiv:2306.08585), "Baseline architecture" → "Determining the code distance": "each
> key generation has a **spacetime volume of 2.6 × 10¹² logical blocks of size d³**. Each such block
> accounts for a d × d patch of physical data qubits operating for d code cycles..."

Dividing by depth alone under-demands the required logical error rate by approximately the logical
qubit count — ~1,400× for Gidney, ~6,000× for Litinski — so the gate passed **earlier** than either
paper supports. That is an optimistic bias, the one direction this project must not drift. The model
now divides by a new `ResourceEstimate.spacetime_volume_logical_qubit_rounds`; where a paper provides
no physical layer, the field is null and `logical_error_budget` reports `undefined` rather than
inheriting an assumption its authors never made.

**ECC-256: 2,126 → 2,330 logical qubits.** `2126` appears nowhere in Roetteler et al.'s LaTeX source.
Their abstract states 2,330, matching Table 2's n=256 row and their own formula
9n + 2⌈log₂n⌉ + 10 = 2,330, identically across arXiv v1–v3. The nearest real figure, **2,124**, is
from a *different* paper — Häner et al. 2020 — which also restates Roetteler's count as 2,338 (+8
qubits for T-depth-1 Toffolis). The likely chain is 2,330 → 2,338 → 2,124 → "2,126". Consequence
worth stating plainly: the site published a wrong number attributed to the wrong paper, and the
correction is a transcription fix, not a change of scientific baseline.

**RSA-2048: 1,409 → 1,537 logical qubits.** 1,409 is the peak *active* count; 1,537 includes idle hot
patches and is the figure Gidney carries into his own physical-qubit and error-budget derivations
("fewer than 1600 logical qubits (including idle hot patches)"). The paper is internally inconsistent
by ~2% across three figures — Table 5 reports 1,399, the stated formula m+3f+2ℓ+len(m) evaluates to
1,432 — and that inconsistency is recorded in the record's `uncertainty_notes` rather than resolved
here. Also recorded: Gidney's 6.5×10⁹ Toffolis are **per factoring, not per shot** (Table 5 caption),
and his 1.25% deviant-shot rate and /0.99 postprocessing factor are *already inside* his 9.2 expected
shots — folding them into `success_probability` as well would double-count repetition.


### ECC-256 default baseline changed to Litinski 2023 — 2026-07-27 (founder decision)

Roetteler et al. 2017 remains the canonical **logical-qubit-count** reference and stays recorded, but
it cannot be the model's default baseline: the paper specifies no code distance, no cycle time and no
runtime, so its spacetime volume is unknown, `logical_error_budget` can only report `undefined`, and
no ECC crossing could ever resolve. Litinski 2023 carries its own physical layer and its own error
budget, which is what the corrected spacetime-volume budget requires.

Figures used are the paper's **2D-local baseline configuration (d=28)**, not its headline
active-volume configuration — active volume assumes a non-local architecture, while the 2D-local
surface-code baseline is comparable with both today's hardware and the RSA-2048 (Gidney, surface code)
baseline. Stated in the paper: 6,000 logical qubits and 109 million Toffoli gates per key at k=1;
484 million logical cycles per key, each logical cycle being d code cycles ("3.8 hours" at a 1 µs code
cycle); a spacetime volume of 2.6×10¹² logical blocks of size d³ per key; and ~20% logical-error
probability after 10 phase-estimation blocks at d=28. Derived here, with the arithmetic recorded in
the record itself: spacetime volume 2.6×10¹² × 28 = **7.28×10¹³ logical-qubit-rounds**, and
success probability (1 − 0.20)^(1/10) = **0.978**.

Title-figure caveat, recorded because it will otherwise be repeated: the paper's title says "50
million Toffoli gates", a number that **appears nowhere in its body** — the body states 109 million
per key at k=1 and 44 million asymptotically.

Unit sanity check (a check, not a validation): the implied requirement of 3.0×10⁻¹⁶ per logical qubit
round is within a factor of ~3 of Gidney's stated 1×10⁻¹⁵ per-round target for RSA-2048.

**RSA-2048 has no `operational` baseline, and that is a finding rather than a gap in curation.**
Gidney 2025's attack takes just under a week, which is the `extended` scenario; the 24-hour
`operational` window it is not. The 2019 Gidney–Ekerå 8-hour attack does fit `operational` but has no
spacetime volume recorded yet. So at the published headline scenario the model has an ECC forecast and
honestly has none for RSA.


## v1 Layer C engineering placeholders (Phase 12, not Part XII decisions)

Deliberately kept separate from "Approved baselines" above: everything in this section is a
documented **engineering interpretation**, the same standing as `forecast.gates`'s
`MINIMUM_OPERATIONAL_UPTIME_FRACTION` constant, not a founder-approved figure traceable to a
citation. See [decision-log.md](decision-log.md)'s "Phase 12 scope" entry for the reasoning.
Stored versioned in `forecast.models.SimulationConfiguration` (not hardcoded), so tuning any of
these is a new configuration version, not a deploy. The first version is seeded from
`forecast.distributions.DEFAULT_DISTRIBUTION_PARAMETERS` by migration `0013` (issue #129), so a
published run is reproducible from the repository alone rather than depending on a configuration
someone created by hand; changing any value below is therefore a new migration. None are calibrated against the Global Risk
Institute survey (Item 7 above) — §17's "Forecast evaluation" reserves that comparison for
Phase 14, so using it as a Phase 12 calibration target would be circular.

| Parameter | v1 default |
|---|---|
| Doubling time — superconducting | lognormal median 2.5y, σ 0.35 |
| Doubling time — trapped ion | lognormal median 3.0y, σ 0.4 |
| Doubling time — neutral atom | lognormal median 2.0y, σ 0.45 |
| Doubling time — photonic | lognormal median 4.0y, σ 0.7 |
| Doubling time — topological / other emerging | lognormal median 5.0y, σ 0.8 |
| Doubling time — modular / networked | lognormal median 4.0y, σ 0.6 |
| Logical-error annual improvement rate (all pathways) | Beta(2, 5) |
| Roadmap on-time probability | 0.3 (global v1, not per-pathway) |
| Roadmap delay if missed | lognormal median 2.0y, σ 0.5 |
| Algorithmic-resource improvement rate (all targets) | Beta(2, 20), capped at 10× cumulative reduction — **the single most consequential, least-grounded placeholder**; deliberately conservative rather than fit to the Gidney 2019→2025 ~20× jump |
| Research disclosure delay (achievement → public knowledge) | lognormal median 0.75y, σ 0.5 — distinct from Phase 13's hidden-capability secrecy-driven disclosure delay |
| Pathway / organisation momentum σ (correlation) | 0.2 / 0.3 |
| Default relative uncertainty (fallback when a capability row's own `*_uncertainty` field is null) | 0.15 |
| Simulation horizon | 60 years |

Only 2 of the 7 threshold-crossing gates (`logical_capacity`, `logical_error_budget`) are
time-varying in v1; the other 5 stay fixed at their currently-recorded values. Growing every
numeric `HardwareCapability` field with an independently-invented rate would be a much larger,
much less defensible invention than growing just the two dimensions this document's Layer A
section already treats as the headline scaling bottlenecks.

**Known current-data limitation**: neither seeded default-baseline `ResourceEstimate`
(`ecc256-roetteler-2017`, `rsa2048-gidney-2025`) sets `success_probability`/`circuit_depth`, so
`logical_error_budget` cannot resolve for either — see decision-log.md's Phase 12 entry.

## Unresolved questions (Part XII)

- ~~Item 5: initial resource-estimate baselines.~~ **Resolved 2026-07-22 — see above.**
- ~~Item 6: initial hardware capability baselines.~~ **Resolved 2026-07-22 — see above.**
- ~~Item 7: initial probability distributions.~~ **Resolved 2026-07-22 — see above.**
- ~~Item 8: whether hidden capability remains scenario-only at launch.~~ **Resolved 2026-07-22:
  yes.**
- Item 9: named scientific reviewer / expert panel / documented review process. **Still
  unresolved — no reviewer named as of 2026-07-22.**
- ~~Item 10: publication day and timezone.~~ **Resolved 2026-07-22: Monday, 00:00 UTC.**
- ~~Item 11: historical hindcasting methodology and minimum evidence archive before public
  launch.~~ **Resolved 2026-07-22 — see above.**

Item 9 may not be silently defaulted by the agent — `IMPLEMENTATION_PLAN.md` Part 0.F item 2.
Forecast-engine production coding does not begin until the "Final implementation gate" list
(end of Part XII) is satisfied — see [decision-log.md](decision-log.md).



---

<!-- evidence-taxonomy.md -->

# Evidence Taxonomy

Version: 1.0 (draft) · Status: draft, pending founder review · Last updated: 2026-07-18

See also: [source-registry.md](source-registry.md), [model-methodology.md](model-methodology.md).
Source of truth: `IMPLEMENTATION_PLAN.md` Part III.

## Core evidence record

Minimum schema fields: `id, evidence_code, title, summary, canonical_url, discovery_url,
source_id, source_type, source_tier, publication_date, retrieved_at, authors, organisations,
doi, arxiv_id, patent_id, language, category, subcategory, architecture,
cryptographic_target, claim_type, raw_claim, normalised_metric_type, normalised_value,
normalised_unit, baseline_value, confidence_source, confidence_verification,
confidence_relevance, confidence_recency, confidence_total, validation_status, reviewer_id,
review_notes, model_inclusion_status, included_model_version, created_at, updated_at,
content_hash`. Full field list: `IMPLEMENTATION_PLAN.md` §9.

## Evidence categories (controlled enum)

`hardware, logical_qubits, error_correction, gate_fidelity, logical_error_rate, runtime,
control_systems, engineering_scalability, quantum_networking, cryptanalytic_resources_ecc,
cryptanalytic_resources_rsa, algorithmic_improvement, compiler_improvement, manufacturing,
funding, patents, roadmap, standard, policy, hidden_capability_assumption, model_correction,
crypto_bitcoin, crypto_ethereum, crypto_wallet_exposure, crypto_protocol_migration,
crypto_public_key_exposure, crypto_attack_runtime, crypto_governance_readiness` (§10).

## Validation statuses

`collected → parsed → duplicate | extraction_failed → under_review → validated |
partially_validated | rejected → scheduled_for_model_run → included_in_model → superseded |
retracted` (§11). Only `validated`, or `partially_validated` with explicit limited weight and
manual approval, may enter the forecast engine.

## Evidence-quality assessment

**Rule**: a multiplicative score is an internal prioritisation heuristic only — never labelled a
calibrated probability or "confidence" without empirical calibration (§12.1). Separate
dimensions tracked independently: `source_reliability, verification_strength, model_relevance,
recency, extraction_quality, independence, measurement_completeness, reproducibility`.

### Public evidence grade

| Grade | Meaning | May change critical model parameters? |
|---|---|---|
| A | Primary technical evidence, strong methodology, independent validation/reproduction | Yes |
| B | Primary technical evidence, disclosed methodology, incomplete independent validation | Yes |
| C | Credible preprint, roadmap, or single-party technical claim | Only after manual approval, in a sensitivity run |
| D | Secondary reporting or incomplete claim | No — discoverable only |
| E | Unsupported, disputed, or unsuitable | No — discoverable only |

Starting-grade guidance (§12.5): independently reproduced peer-reviewed result → A; peer-reviewed
with disclosed methods, no reproduction → B; high-quality technical preprint → B or C after
review; company technical paper with sufficient measurement detail → B or C; company roadmap →
C; company press release → C or D; news report linking to primary evidence → discovery-only D;
unsupported social-media claim → E. These are governance defaults, not universal constants.

### Internal prioritisation score

May be calculated for review triage, but: keep components visible, cap recency's influence, add
an independence penalty for duplicated claims, never let popularity substitute for technical
quality, version the scoring configuration, and test whether it actually predicts reviewer
decisions before relying on it (§12.3).

### Retractions and corrections

Use DOI/Crossref metadata and publisher notices to detect corrections, retractions, expressions
of concern, updated versions, or a preprint superseded by journal publication. A material
retraction/correction triggers: evidence-status update → identify affected model versions → new
model run where impact is material → visible correction note in model history (§12.4).

## Unresolved questions (Part XII)

- Item 12: whether evidence summaries are written manually or AI-assisted with mandatory human
  review.



---

<!-- hidden-capability-model.md -->

# Hidden Capability Sensitivity Model

Version: 1.2 · Status: item 8 founder-approved (2026-07-22); Phase 13 implemented (2026-07-23) ·
Last updated: 2026-07-23

See also: [model-methodology.md](model-methodology.md), [event-definition.md](event-definition.md).
Source of truth: `IMPLEMENTATION_PLAN.md` Part IV, section 15, and Part II section 8.

## What this is not

A military/intelligence/undisclosed-private-programme adjustment factor is never represented by
fabricated news or unsupported claims, and it is not a data feed — see
[source-registry.md](source-registry.md), §"Hidden capability is not a data feed". It is
implemented as a documented, versioned prior distribution over scenarios: no meaningful lead,
limited lead, moderate lead, major lead, extreme low-probability lead.

## Do not blend arbitrary probabilities into the main forecast

An example lead-time mixture (0–1y / 1–3y / 3–6y / 6–10y) is a placeholder only, and stays a
placeholder until supported by formal expert elicitation and governance approval — it must never
silently become production configuration.

## Initial public implementation

Publish four bands, never merged into the public-evidence baseline:

- `Public-Evidence Baseline`
- `Hidden-Capability Sensitivity: low`
- `Hidden-Capability Sensitivity: moderate`
- `Hidden-Capability Sensitivity: severe`

For each sensitivity scenario, disclose: assumed lead-time distribution, hidden
algorithmic-improvement assumption, engineering-advantage assumption, disclosure-delay
assumption, and the resulting P10/P50/P90 movement versus the public-evidence baseline.

## Governance

Hidden-capability settings: cannot be altered by ordinary evidence-feed records; require
two-person approval once the project has two qualified reviewers; require a written audit
reason; create a new configuration version on every change; must be visible in the public
methodology pages; must never be described as known classified capability.

## Required public wording

> No classified capability is claimed as known. These scenarios show how undisclosed progress
> could change the public-evidence baseline.

## Resolved questions (Part XII)

- Item 8: **resolved 2026-07-22 — yes, hidden capability remains scenario-only at launch.**
  Founder-approved as drafted above: never blended into the public-evidence baseline, published
  as four separate bands. See [decision-log.md](decision-log.md).

## Implementation (Phase 13, §43)

`HiddenCapabilityConfiguration` (`forecast/models.py`) is a versioned, admin-managed model with a
governance shape not used anywhere else in this codebase: creating a row only **proposes** it
(`proposed_by`/`proposed_at`); it only becomes usable once a **different** user runs the
`approve_configuration` admin action (`forecast/admin.py`), setting `approved_by`/`approved_at`.
`get_active_hidden_capability_configuration()` (`forecast/models.py`) returns the highest version
with `approved_by` set, or `None` — an un-approved proposal is never active. With only one admin
account today (item 9, named reviewer, still open — see below), nothing can be approved; this is
the intentional, documented consequence of the two-person rule above, not a bug.

Each of the three sensitivity levels (`low`/`moderate`/`severe`) is stored in one JSONField,
`sensitivity_scenarios`, as:

```json
{
  "lead_time_mixture": {
    "distribution": "mixture",
    "components": [
      {"scenario_label": "no_meaningful_lead", "weight": 0.8, "distribution": {"distribution": "fixed", "value": 0.0}},
      {"scenario_label": "limited_lead", "weight": 0.15, "distribution": {"distribution": "lognormal", "median": 1.0, "sigma": 0.3}},
      "... moderate_lead, major_lead, extreme_low_probability_lead ..."
    ]
  },
  "hidden_algorithmic_improvement": {"distribution": "lognormal", "median": 1.05, "sigma": 0.1},
  "engineering_advantage": {"distribution": "lognormal", "median": 1.05, "sigma": 0.1},
  "disclosure_delay": {"distribution": "lognormal", "median": 2.0, "sigma": 0.5}
}
```

The 5 canonical `scenario_label` values (§8's list) are required exactly once each per level;
weights must sum to 1.0 (a new generic `"mixture"` kind in `forecast/distributions.py`'s
distribution registry, reusable outside this module). `forecast/hidden_capability.py` holds the
canonical labels/descriptions, the `SENSITIVITY_LEVELS` tuple, the required public disclaimer
string (verbatim above), and `apply_hidden_capability_adjustment()` — the pure function that
turns a sampled lead-time-in-years plus two multiplicative discount factors
(`hidden_algorithmic_improvement`, `engineering_advantage`, each centered near 1.0) into a single
effective lead-time, subtracted from the public-evidence-baseline crossing date.

**This composition (`lead_years × algorithmic_factor × engineering_factor`, plus the
`DEFAULT_HIDDEN_CAPABILITY_PARAMETERS` numeric placeholders themselves) is an unverified v1
engineering assumption, not a Part XII decision or a scientifically calibrated model** — the same
caveat class as `forecast.gates.MINIMUM_OPERATIONAL_UPTIME_FRACTION` and
`forecast.distributions.DEFAULT_DISTRIBUTION_PARAMETERS`. No Part XII item covers the actual
hidden-capability parameter *values*, only item 8's scope question above. See
[decision-log.md](decision-log.md)'s Phase 13 entry.

**Known v1 limitation, documented not silently fixed**: a scenario that never crosses within the
horizon under the public-evidence baseline stays censored under every hidden-capability band too
— the adjustment only pulls an already-crossing scenario earlier, since honestly deciding whether
a hidden lead turns a non-crossing scenario into a crossing one would require re-running the full
time-varying gate search under boosted capability, out of scope for this phase.

`forecast.simulation.run_simulation()`'s new `include_hidden_capability: bool = True` parameter is
§43 item 7's exclude switch; `SimulationRun.hidden_capability_report` (JSONField, never blended
into the `ecc_public_p50` etc. fields on the same row) always carries an honest
`excluded_reason` (`"switched_off"` / `"no_approved_configuration"` / `"configuration_not_
approved"`) instead of a comparison when one wasn't computed. The last of those three closes
issue #39: an explicitly-passed `hidden_capability_configuration` (e.g. via `run_simulation`'s
`--hidden-capability-configuration-version` flag) is rejected in `_run_hidden_capability_
scenarios()` itself when `approved_by` isn't set, the same §15.3 two-person gate
`get_active_hidden_capability_configuration()` already enforces for the default selection —
closed at the one function every caller funnels through, not just filtered in the CLI. `run_
simulation` management command flags: `--exclude-hidden-capability`,
`--hidden-capability-configuration-version` (raises `CommandError` up front for an unapproved
explicit version, rather than silently degrading to a report field).
