εCryptanalytic Effort
technical companion · working draft · 9 October 2026

How the numbers
are made

Definitions, assumptions, implementation details and unresolved problems

Back to the readable summary · White paper (PDF)

Purpose of this page

This is the audit trail behind the summary. It describes what the code actually does, where the inputs come from, and where the interpretation goes beyond the evidence. It preserves known discrepancies rather than presenting an idealized method. The data and code links below are pinned to repository snapshot 2525642. The corpus and labor sections document that snapshot; they do not report a new classification or re-estimation run. The algorithmic-history section is a later, separately sourced addition.

What is being measured?

The study starts with records of publications and reported computations. Almost everything called an effort, year or budget downstream is a transformation of those records. The distinction matters: precision in a table does not make the underlying quantity directly observed.

The chain from evidence to interpretation.
QuantityHow obtainedWhat it does not establish
Publication metadataBibliographic harvests and abstracts.Completeness, correctness, or time spent on the work.
Inclusion and problem weightsModel classifications against a rubric.Independent expert endorsement of the labels or papers.
Fractional research itemsIncluded items allocated across assumption classes.Equal difficulty, value or labor per item.
Research-allocation yearsPublication shares over career windows, plus a paper-prior fallback.Actual full-time staffing or recorded working hours.
Nominal person-hoursYears multiplied by 2,000 hours/year.A time-use measurement.
Equivalent inference budgetsAssumed human cost divided by assumed AI productivity.A fitted exchange rate, attack forecast or measured labor replacement.
Computational frontier gapsRatios under a selected attack-cost model.The research effort needed to discover a better attack.

The three principal classes are ECDLP, lattices and syndrome decoding. Factoring/finite-field discrete logarithms, isogenies and multivariate systems provide comparators. The labels collect broad mathematical families, not just ordinary instances of current encryption standards. More published scrutiny is not a numerical security guarantee.

Acquisition, screening and deduplication

Sources and stages

The initial ePrint harvest parses 28,087 identifiers and retains 27,249 live records. The multi-source extension parses about 8.7 million DBLP records, retaining 375,911 after a title/venue/thesis screen. OpenAlex contributes 50 phrase searches (14,195 works), abstract enrichment for 39,207 core-venue DBLP DOIs (roughly 19,000 abstracts found), and citation expansion from 811 resolved ePrint seeds. The latter returns 5,071 backward references and 16,216 forward citers, with a cap of 400 citers per seed. Forty-six zbMATH queries return 7,498 rows. These are overlapping acquisition-stage counts; adding them would not produce a count of distinct papers.

A direct arXiv API acquisition was abandoned after repeated rate-limit/server failures. arXiv works indexed by other sources can still enter the corpus. Older publications, items without useful metadata and work outside these databases remain potential gaps. Raw harvest caches are ignored by Git and are not part of the published snapshot.

The merged acquisition contains 356,673 items; the first keyword screen retains 19,266, including 1,875 already covered by ePrint candidates. Tier A prioritizes promising keyword signatures; tier B contains 8,353 lower-priority candidates. A second acquisition/classification round adds 1,840 candidates from a broader screen and a seeded random sample of 835 tier-B items. The combined classification outputs contain 16,471 rows before three documented manual exclusions, leaving 16,468 labeled rows.

What “same paper” means

The source merge uses DOI matching first and normalized-title matching otherwise. Its title key lowercases, removes everything outside ASCII letters and digits, and uses the first 60 characters; titles with fewer than ten normalized characters are dropped in that merge. Metadata merging prefers the earliest year and longest abstract; some fields, including authors and venue, use the first nonempty value. This is bibliographic deduplication, not a full reconciliation of a work's preprint, conference, journal, translated and thesis versions.

Changed titles and translations can survive as duplicates. Shared title prefixes can merge distinct works. Non-Latin titles can be poorly represented. A long abstract need not be the correct abstract. DBLP's early position in acquisition affects which author/venue metadata survive. These decisions propagate into both counts and identities.

Inspect: acquisition report, source merge, screening rules.

What gets classified as cryptanalytic work?

Classifiers see title, authors, year, available venue/category metadata and abstract, rather than routinely reading full papers. Later first-pass batches cap abstracts at 1,000 characters; the Phase 1 second pass uses up to 1,500 (2,200 in the earlier pilot). Labels contain inclusion, category, weighted problem tags, target parameters, layer, lattice regime, leakage, quantum/record flags, a confidence score and a short reason.

Category rules in the main cryptanalysis rubric.
CodeMeaningMain effort view?
AClassical algorithms, including faster implementations of mathematical solvers.Yes
BQuantum algorithms or quantum resource estimates.Yes
CStructural results, special instances, weak keys, distinguishers and structural attacks.Yes
DConcrete cost models, security estimates and parameter analysis.Yes
EComputational records with resources reported.Yes
FNegative results or limits on a class of attacks.Yes
HSerious claimed attacks later withdrawn/refuted, or their refutations.Yes
GHardness reductions or equivalences without an attack.No; tracked separately
SSurveys, tutorials, SoKs and theses centered on attacks.No; tracked separately
XOut of scope.No

Inclusion is a substantive choice

yes means the main contribution fits the rubric; borderline means scope or relevance is uncertain; no excludes the work. The main class totals use only yes and categories A–F/H. They exclude construction papers that merely use an assumption, implementation or protocol flaws that say nothing about the mathematical problem, and extraction of physical leakage through timing, power analysis or faults.

They deliberately include mathematical attacks given abstract hints or partial information, speed improvements to lattice reduction and other solvers, special/weak instances and quantum analyses. A lattice attack on an abstract leakage model is different from the engineering needed to extract that leakage. Neither necessarily attacks an ordinary deployed instance. FHE parameter regimes and signature-motivated shared machinery can also contribute.

The main rubric explicitly excludes HAWK/lattice-isomorphism as a distinct assumption, although the separate HAWK case study intentionally includes it. Scope rules evolved during classification: additions concerned CSIDH, low-genus Jacobian DLP, sensational unvetted claims and generic lattice-reduction algorithms. Some ambiguous boundaries remain, including nonce-leakage/hidden-number problems and non-Hamming decoding.

The classifier's confidence is its self-assessment, not a calibrated probability. Source restriction and a confidence ≥0.6 sensitivity view do not establish mathematical correctness. A documented manual exclusion removes a particular 2026 “Module Lattice Security” series; it is not evidence of systematic human validation of the rest.

Inspect: full rubric and examples, manual exclusion ledger, bounded scope audit.

Attribution, denominators and source views

For paper \(i\), normalize its supplied tag weights \(t_{ij}\) to sum to one, then apply the fixed tag-to-class map \(M_{jc}\). The resulting class weight is:

$$w_{ic}=\sum_j\frac{t_{ij}}{\sum_k t_{ik}}M_{jc},\qquad P_c=\sum_{i\in I}w_{ic}.$$

\(I\) is the selected inclusion/category set. Fractional counts add \(w_{ic}\); touch counts count every selected item with nonzero class weight. A paper can touch several classes. Author-name counts in the paper reports are different from the subsequently resolved person identities.

The implemented attribution map, including unallocated mass.
TagsHeadline allocation
ecdlp100% ECDLP, including the rubric's low-genus generalization.
factoring, ffdlp100% combined factoring/finite-field DLP.
lattice_generic, sis50% structured lattices, 50% plain LWE.
lattice_structured, ntru100% structured lattices.
lwe_plain100% plain LWE.
sd_generic, sd_goppa, sd_qc100% syndrome decoding.
lpn50% syndrome decoding; the other half is unallocated.
isogeny, multivariate100% to their respective comparator classes.
rank, otherNo headline allocation.
rsa_variantPresent in the rubric but absent from the map: no headline allocation in this snapshot.

Combined lattices add the structured and plain weights; their union touch count counts a shared paper once. Summing combined lattices and its two subrows would double-count. The 50/50 split is a convention, not a measured division of research benefit.

Why the totals do not all add up

The export has 4,349 records = 3,520 yes + 829 borderline. Of the yes records, 2,877 are in the main effort categories, 333 are G and 310 are S. The 2,877 effort items allocate only about 2,698.9 fractional items across the seven disjoint headline classes; 134 have no headline weight. Missing mappings and deliberately partial allocations explain why “items in the corpus” is not the sum of the class rows.

The website's annual paper series cover 1970–2026. They omit an undated ECDLP item and Prange's 1962 decoding paper, each with class weight one. The summary table uses full-corpus totals and retains them. Annual values are rounded by the data builder, so summing chart values can also differ slightly from directly summing unrounded records.

Alternate views are alternate assumptions

The main site counts observed tier-B sample papers once. A separate sensitivity view weights those sampled items by \(8353/835\approx10.004\) to represent the full tier-B population. That is a sampling extrapolation, not ten real papers for every observed one. Only 26 sampled items are yes (another 11 are borderline), so the extrapolation is sensitive to a small number of labels. Author counts are not inflated by this factor.

The site's “ePrint / DBLP” subset uses the export's source flag. It is not a peer-review or verification filter: ePrint is a preprint archive and DBLP also indexes informal work. There is a known legacy inconsistency: the older counting script recognizes round-two ePrint-only source lists, while the export also accepts mixed-source lists containing ePrint. Ten main-view records differ. This adds 1.5 ECDLP, 1.0 combined lattice and 7.5 factoring/FF-DLP fractional items to the website subset relative to the older report. All-sources headline counts are unaffected.

The often-quoted lattice/ECDLP span of roughly 1.3–2.6 mixes author counts, paper counts, source restrictions and sampling assumptions. Its endpoints are not uncertainty quantiles for one well-defined estimator.

Inspect: attribution map, counting implementation, export rules, record-level export.

What the quality checks establish

Two model passes are not two independent judgments

The second pass reviews first-pass yes/borderline items plus approximately 10% of first-pass exclusions. It can see the first label. Reported inclusion agreement of roughly \(\kappa=0.84\text{–}0.89\) is agreement on the selected reviewed subset, not a representative estimate of accuracy against expert truth. Correlated model errors, rubric mistakes and omitted full-text evidence can survive both passes. Where a second-pass label exists, it replaces the first; unreviewed exclusions retain the first-pass label.

Some repository reports call exclusion-audit reversals a “false-negative rate.” The actual denominator is sampled first-pass exclusions. Round one reverses 7 of 648 audited exclusions (1.08%); round two reverses 1 of 213 (0.47%). These estimate the fraction of rejected candidates the second model would reinstate under the sampling assumptions. They are not the fraction of all relevant papers missed, and subtracting them from 100% does not give classifier recall.

The reference-list benchmark

Twelve seed papers supply 429 reference-list entries, not necessarily 429 distinct papers. Of these, 421 were harvested, 329 screened into the pool, 327 labeled and 295 labeled yes/borderline. One code-survey seed supplied no references through the API. Some uncaptured or excluded references are properly out of scope.

The approximately 96–100% figures in the narrative reports are judgments about in-scope coverage after inspecting gaps in these selected lists. They are not \(295/429\), a field-wide recall measurement or a held-out test: the benchmark itself helped motivate the broader second-round screen. Source bias and shared citation neighborhoods remain. The documented systematic human spot-check of corpus labels has not been completed.

Inspect: second-pass prompts, agreement calculation, round-one audit, round-two audit, reference-list benchmark.

From author names to research-allocation years

Identity and publication histories

Items are matched to DBLP by record key, ePrint identifier or DOI. Normalized-title matches also require years within ±1 when both are available. Candidate matches prioritize overlap with the item's author list. DBLP person aliases are merged; homonym-numbered identities remain distinct. Authors lacking a direct match are tried against a unique full name and then a surname/initial match within the cryptanalysis population, with first-name compatibility checks. Unresolved names remain separate string identities.

The author map reports 7,613 authorships on main-view items and approximately 3,240 identities. About 2,573 have usable DBLP/zbMATH publication histories; the remaining 667 use a fallback. These are estimates of people, not a verified census. Alias failures can split a person, and an incorrect match can merge different people.

Career output merges DBLP publications and accepted zbMATH author profiles by normalized-title hash. A zbMATH match uses identifiers or overlapping known work, not name similarity alone. The goal is to avoid allocating all of a mathematician's research to cryptanalysis just because DBLP misses their mathematics publications. Publication history still misses unindexed papers, other duties, unpublished research and nonpublishing contributors.

The implemented window formula

For person \(p\) and year \(y\), let \(I_{py}\) contain their included effort items dated within \(y-2\) through \(y+2\). Let \(B_{py}\) be all distinct career publications in that window. Let \(v_{py}\) be the number of window years between the first observed career publication and the 2026 observation boundary. The default rate floor is \(r=1\) publication per available year. The actual denominator and class allocation are:

$$d_{py}=\max\left(B_{py},\sum_{i\in I_{py}}\max_c w_{ic},r v_{py},10^{-9}\right),\qquad a_{pcy}=\frac{\sum_{i\in I_{py}}w_{ic}}{d_{py}}.$$

The inner maximum is over disjoint headline classes, not the additional combined-lattice row. Contributions are summed across person-years; the per-paper fallback described below is added. A person counts from their first observed career year to their last. Someone whose last career publication is in 2025 or 2026 is carried through 2026 to allow for indexing lag. The window can include future publications relative to \(y\), so this is a retrospective allocation, not an observation of what researchers were doing at that moment.

Each coauthor receives the paper's class weight in their own publication-share numerator. The numerator is not divided by author count. Division by authors is used for a separate publication-credit statistic and the missing-career fallback. Equal publication credit is not a measurement of equal labor contributions.

For an interior five-year window with two topic papers and ten total papers, allocation to the topic is \(2/10=0.2\) in that year. One topic paper and one total indexed paper in that window gives \(1/5=0.2\), because the one-publication/year floor dominates. Five topic papers and five total papers can give one full allocation year; it says nothing about whether the person actually spent a full working year on research.

A known implementation weakness

The denominator's sum(max(class weights)) term does not guarantee that total allocation across disjoint classes is at most one. A paper split 0.5/0.5 contributes only 0.5 to this guard while contributing one across its class numerators. Observed publication totals and the rate floor may mask the issue. A sum of each paper's disjoint class weights would express the intended aggregate cap more directly.

This defect is documented, not silently repaired in the published figures. Its current numerical impact has not been established: the ignored raw DBLP match/profile caches needed for an exact rerun are absent from this checkout. The baseline exports remain the source for the site and paper.

Inspect: identity matching, career enrichment, implemented allocation formula, method notes.

Fallback priors, career adjustment and hours

The author estimate is a hybrid

For a person with no usable career history, each included paper contributes its class weight times a category effort prior, divided by the paper's resolved number of authors. The main fallback midpoints are 1.5 total person-years for A/B/C/F, 1.0 for D, 6.0 for E and 0.75 for H. These are assumed total years across the author team, not years per author. Only the missing-career authors' shares enter this fallback.

Dated fallback effort is spread uniformly across the same ±2-year neighborhood, truncated at 2026. An undated item can enter the per-person total but cannot enter the dated series. Consequently the person table and cumulative chart need not reconcile exactly. In this snapshot, an undated ECDLP item contributes 1.5 fallback years to the former but not the latter, alongside rounding differences. Cumulative labor totals also retain allocations from before 1970 even though the exported annual rows start in 1970: syndrome decoding's first cumulative value is 0.5 years versus 0.2 in that year's annual row.

Fallback priors supply approximately 16.9% of the dated ECDLP total, 9.0% of lattices and 13.1% of syndrome decoding. Therefore apparent agreement between a paper-prior estimate and the author estimate is not an independent validation: the latter partly contains the former, and both share the same literature labels.

Career-stage fractions

The adjusted view multiplies publication-share contributions by assumed fractions of working time spent on research. It leaves the paper-prior fallback unchanged, treating those priors as already expressed in person-years.

Research-time fractions used by the adjusted model.
StageAssumed fraction
Before the PhD year0.90
PhD year and following four years0.85
Established, academic sector0.45
Established, nonacademic sector0.70
Established, unknown sector0.50

PhD dates prefer Math Genealogy information, then ORCID doctoral entries, then an estimate from first publication and an observed median offset of two years (564 usable offsets, interquartile range zero to four years). Employment sector comes from dated ORCID organization-name heuristics, otherwise “unknown.” About 76.9% of the PhD dates among included people with output profiles are estimated. These fractions and stages are conventions, not time-use data.

Cumulative dated totals used by the website, through its 2026 observation window.
ClassAllocation yearsAdjusted years
ECDLP593.8438.9
All lattices1,036.4758.5
Syndrome decoding331.4248.7
Factoring / FF-DLP1,111.1847.5
Isogenies116.685.2
Multivariate210.2149.7

Why the totals may be too high

Publishing only on a topic does not mean working on it full time. Teaching, administration, consulting, other employment, short projects and unequal coauthor contributions can make publication share substantially exceed time share. Missing nontopic publications can inflate that share further. Career adjustment lowers the principal totals by roughly a quarter but does not validate them. There is no measured basis for a tight numerical error bar or a claim that the adjusted series is the true answer.

Other biases can operate in the opposite direction: prolonged failed projects, nonpublishing engineers, private research and work before a first publication are poorly represented. Opposing biases do not establish cancellation. Missing classified work does not turn an uncertain public-time proxy into a guaranteed lower bound.

The paper uses \(2{,}000\) hours per research-year to express nominal person-hours. At that convention, ECDLP's 593.8 allocation years become 1,187,600 nominal hours. Using 1,500 or 2,200 hours/year scales all hours by 0.75 or 1.10; it does not change the year-based dollar scenarios. None of these conversions reveals logged hours.

Inspect: annual and cumulative export, person-level export, career-stage acquisition.

Effort around HAWK and mathematical problems

The named-problem study uses curated seeds, citation expansion and a separate rubric. It estimates work on a problem or related approaches, rather than counting papers that merely assume a conjecture. Its scopes include HAWK/module-LIP, Unique Games and related lines, 2-to-1 Games, and 22 selected Erdős problems. Overlapping literatures such as Unique Games and 2-to-1 must not be added as independent totals.

For item \(i\), relevance \(u_i\in[0,1]\) is supplied by the classifier. Its inclusion multiplier \(b_i\) is one for yes and one-half for borderline. For category effort prior \(\mu_{k_i}\), the estimator is:

$$H_P^{\mathrm{papers}}=\sum_i b_i u_i\mu_{k_i}.$$
Problem-study priors differ from the author-map fallback.
CategoryProblem-study prior, total years/paperMain author fallback midpoint
A, C, F1–2; midpoint 1.51.5
BNot part of this calibration rubric1.5
D0.5–1.5; midpoint 1.01.0
E2–10; midpoint 6.06.0
H1–2; midpoint 1.50.75
S0.5–1; midpoint 0.75Not in main effort view
GExcluded despite being labeledNot in main effort view

Displayed central values sum midpoint priors. The chart's ranges sum all lower and upper endpoints; they are sensitivity ranges, not confidence intervals. Independent random draws per paper can produce a narrow interval without resolving a shared error in effort per paper. HAWK's older 33.2–37.6-year Monte Carlo interval around 35.4 years omits that shared uncertainty, missing work and attribution error.

Human labor and chronological cutoffs are not audited

The site preserves the paper-comparable literature totals and flags explicitly AI-assisted entries. It does not automatically remove them or impose a pre-discovery publication cutoff. The flag depends on disclosure visible in the abstract; an unflagged item is not proof of purely human authorship. Erdős #741's entire 0.495-year estimate is flagged AI-assisted; for #846, 1.5 of 1.875 years is flagged. Those totals cannot honestly be called prior human effort. The website labels them literature-effort scenarios.

“Solved problem” can hide variant and part distinctions. The sample retains studies of partial resolutions and related statements as well as complete resolutions, and the website does not independently verify the underlying mathematics. Counts of manuscripts, conjectures, families, attempts and formal statements are not interchangeable.

An old question may have received little effort

A problem can remain open for decades because few people seriously tried to solve that particular question. A cheap AI solution can uncover an overlooked opportunity rather than overcome sustained human resistance. Related papers may have devoted little effort to the exact question, while unsuccessful attempts may leave no publication. These selected successes are not a representative sample of mathematics, and cannot calibrate the cost of replacing a mature field's cryptanalytic scrutiny.

Inspect: problem rubric, item-level problem labels, website estimator and flags.

Economic scenarios, tokens and disclosure boundaries

The multiplier is assumed

The main scenarios use a loaded human cost convention \(c=\$250{,}000\) per research-year and assumed economic productivity \(s\in\{1,10,100\}\). Larger \(s\) means more comparable research per inference dollar. No scenario is a preferred estimate or an uncertainty percentile. The explorer opens at 10× only as an illustrative setting.

$$D_X(s)=\frac{cH_X^{\mathrm{share}}}{s},\qquad H^{\mathrm{equiv}}(D;s)=\frac{sD}{c}.$$

One estimated human year becomes $250,000, $25,000 or $2,500 of inference. A $100,000 AI budget becomes 0.4, 4 or 40 assigned human-year equivalents. Applying the same multiplier across classes preserves their ranking by construction; it does not show that AI transfers equally well to every class or to unsuccessful lines of inquiry.

These are inference allowances. Human supervision, verification, tool compute, infrastructure and model training are outside the allowance; historical human attack-computation costs are likewise not included in \(cH\). A public API-price valuation is not an internal compute bill. If the human-effort estimate is too high by a factor of two, its equivalent budget is too high by the same factor.

Token equivalents require a price and mix

$$p_{\mathrm{mix}}=f_i p_i+f_c p_c+f_o p_o,\qquad T_{\mathrm{ref}}=\frac{10^6D}{p_{\mathrm{mix}}},\qquad f_i+f_c+f_o=1.$$

The fractions describe uncached input, cached input and output tokens, priced per million tokens. The site's default $10/million is hypothetical. A real comparison must specify model version, date, caching, token mix and billed reasoning tokens. Dollars alone do not identify an actual token count; elapsed time and concurrent agents do not identify tokens or human hours.

HAWK normalization is a separate assumption

$$D_X^{\mathrm{HAWK}}=\$100{,}000\frac{H_X^{\mathrm{papers}}}{H_{\mathrm{HAWK}}^{\mathrm{papers}}}.$$

HAWK has 35.4 estimated paper-based years using yes-only work, or 38.4 with borderline work at half weight. These imply roughly $2,825 or $2,604 per historical paper-based year if one assigns the reported AI discovery budget to the preceding literature. That transfer is assumed, not measured. Both sides must use the same paper estimator; a shared prior scale can then cancel, but differences in coverage, relevance and category mix do not.

The site removed its former HAWK preset because it applied this paper-derived rate to author-output years. The remaining $100,000 marker is a disclosed discovery budget only. The human team's active hours are undisclosed; the older 0.3–1-year suggestion was hypothetical. Human and AI HAWK results also achieve different dimension reductions. Publication dates do not establish independent discovery or duration; Anthropic reports private disclosure to the HAWK team in June.

What the cost markers count

Selected disclosures retained with their accounting boundaries.
CaseDisclosed figureBoundary
Anthropic HAWKAbout $100,000 API cost; about 60 hours.Discovery process, not attack execution; tokens and active human hours unspecified.
Erdős #1$405 and $1,384; mean $894.50.Two priced successful runs, including one excluded from the fixed-budget score. Unsuccessful runs omitted from this mean.
Erdős #74 and #126Six successes averaging $158.83; four averaging $211.Means over disclosed successful attempts of the prerelease Astra model, including additional non-systematic attempts.
Erdős #548 and #571$363 and $617.One priced successful attempt each; not total cost to obtain each result.
OpenAI ten advancesRoughly $2,000 at Sol API rates.Reference-price valuation of solution tokens across ten results; not a complete campaign bill.
OpenAI Navier–Stokes claimAbout 130 billion output tokens for the result; about 300 billion across the wider campaign.The smaller number is a subset, not additive. Input tokens and an all-in dollar bill are undisclosed.

The Erdős CSV records individual successes and unpriced failures. The builder averages only priced outcomes marked resolved (including resolved but excluded from the score), restricted to the named prerelease Astra model. It does not divide total campaign spend by all attempts or all distinct results. The primary paper reports more than $220,000 across the wider campaign. Its roughly $10,000 “expected cost” discussion is a budget/success-rate heuristic, not a reconciled campaign ledger, and is not used as a calibrated reference line.

For the Navier–Stokes output-only example, a hypothetical output price \(p_o\) dollars/million gives \(D=130{,}000p_o\) and \(H^{\mathrm{equiv}}=0.52sp_o\). This prices only disclosed output tokens; it does not estimate human historical effort on Navier–Stokes or the internal system's actual cost.

Inspect: scenario definitions, disclosure ledger and primary links, attempt-level costs, Erdős primary report, HAWK primary report.

Computational records and frontier gaps

The separate record track has 338 entries; ten were marked user-verified in the snapshot, including records accepted provisionally. The remainder are agent-transcribed. A record's native problem size, algorithm, structure exploited, hardware and time belong together. A special-structure solve or restricted interval discrete logarithm is not a generic solve at the same nominal bit length.

The frontier statistic is \(g=\log_2 W(\text{target})-\log_2 W(\text{record})\). It is a ratio within one model, not the target's security strength. A common size-independent multiplicative factor cancels; changing hardware efficiency, memory cost and size-dependent corrections do not. Cross-family gaps are not equally calibrated dollar or runtime distances.

Models behind the website's selected frontier rows.
TargetModel and anchorMeaning of the gap
P-256Generic Pollard rho, approximately \(0.886\sqrt{\ell}\) group operations; record subgroup about 117.3 bits.About 69.3 bits of modeled group-operation work. Costs per operation differ between curves.
ML-KEM-512 / 768Sieving dimension 171 in a dimension-210 approximate-SVP record, versus target effective sieve dimensions 375 / 586.60–75 / 121–152 bits per sieve call under slopes 0.292–0.367. Not complete attack-to-record runtime ratios.
Classic McEliece 348864Goppa decoding record \(n=1409\), about \(2^{63}\) on the challenge's scale; proxy target \(n=3471\).About 79 bits on that scale. Proxy parameters differ from the deployed instance.
HQC-128QC decoding record \(n=4098\); Prange iteration estimates and a seven-bit quasi-cyclic discount.About 61 bits under a coarse same-model approximation.
RSA-2048Leading GNFS expression at 896 and 2048 bits.About 35.02 bits, or 35 billion times modeled work; omitted asymptotic terms are uncalibrated.

Why the RSA gap is only about 35 bits

For a \(b\)-bit modulus, substituting \(N\approx2^b\) into the leading GNFS expression gives the index:

$$F(b)=\frac{(64/9)^{1/3}(b\ln2)^{1/3}[\ln(b\ln2)]^{2/3}}{\ln2}.$$

\(F(896)=81.8596\), \(F(2048)=116.8838\), and their difference is 35.0242. GNFS scales subexponentially with modulus size. These values are not measured operation counts. NIST's separate 112-bit category for RSA-2048 is not used as an endpoint; subtracting 81.86 from 112 would mix conventions. The decimal places document arithmetic, not forecasting precision.

Lattice and hardware caveats

The SVP challenge accepts a norm within 1.05 times the Gaussian heuristic. The modeled frontier uses the solver's reported maximum sieving dimension, not its ambient lattice dimension. Target dimensions come from a selected primal-attack estimate with dimensions-for-free conventions. The slope interval contrasts an asymptotic value with a GPU-sieving fit; it is not a statistical confidence interval. The project's own smaller fitted slope is not used because chronological records confound size with changing algorithms and hardware.

Repeated sieve calls in the full ML-KEM-512 target attack contribute roughly nine additional modeled bits. The record solve also has work outside a single sieve call, so adding nine is not by itself a complete normalization of both whole computations. Memory, parallelism and algorithm selection require separate treatment.

Hardware normalization is another assumption layer

The record report converts reported compute into ranges of “2020 core-year equivalents.” Its parser uses the first available route in this order: stated core-years times a CPU-era factor; stated GPU-years times a workload factor; parsed MIPS-years; parsed CPU-hours divided by 8,766 and era-adjusted; or a parsed device/core count times wall-clock days divided by 365.25. If a run takes less than two days and no hardware count is parsed, it assumes one machine. Otherwise the value remains missing. A parser taking the wrong number or unit can therefore matter.

Configured conversion factors; ranges are assumptions, not confidence intervals.
Input unit2020 core-year equivalents
CPU core-year, by era1990: 0.01–0.03; 1995: 0.03–0.06; 2000: 0.08–0.15; 2005: 0.25–0.40; 2010: 0.50–0.70; 2015: 0.80–0.90; 2020 onward: 1.
GPU-year, by workloadLattice: 100–500; ECDLP: 50–200; decoding: 20–100; factoring: 300–600; other: 20–500.
PS3-year / FPGA-year for ECDLP3–8 / 2–10.
MIPS-year0.00002–0.00005.

CPU factors use the latest configured era no later than the record year; an unknown or earlier year uses the earliest entry. Missing device/workload combinations fall back to the broad “GPU: other” range, even for a differently named device. These factors cannot reconcile memory, communication and algorithm-specific hardware efficiency. They estimate computation, not researcher years; missing compute does not mean zero effort.

“Record-setting” here means a strict increase in numeric instance size, ordered by date within groups defined by class, problem, challenge and structure flag, with additional LWE noise-parameter groups. Each low-weight decoding entry is its own group. This is a mechanical rule, not an expert ranking of all instances' difficulty. Totals add only record-setting rows with usable compute, with Bitcoin puzzles reported separately; the export retains other rows for inspection. The lattice frontier is restricted to the Darmstadt SVP challenge, and the generic factoring frontier excludes structured/SNFS records.

Inspect: full normalization method, parameters, source ledger, normalization code, individual records.

Selected algorithmic bounds and attack events

This addition documents the revised algorithmic-history track. Unlike the corpus and labor links above, its source links point to the current history ledger. It is a deliberately selected history of supported analyses, not a complete reconstruction of historical beliefs, a measured security-loss curve or a forecast of the next break.

Numerical indices and their inverses

For a track with reference size \(s_0\), each eligible analysis supplies a leading-term function \(I_a(s)\). At year \(t\), the plotted index is the minimum among selected analyses dated by then:

$$I_t(s_0)=\min_{a\in A_t} I_a(s_0),\qquad s_t(128)=\max_{a\in A_t} I_a^{-1}(128).$$

The inverse view independently takes the largest threshold needed to make every selected attack index at least 128. Algorithms can cross, so the two views need not use the same winning analysis. Same-year points show the year-end envelope; intermediate results remain in the ledger. Relative sizes divide by the latest threshold within that track. They are not key-size recommendations. Omitted factors can differ across algorithms, so even a within-track difference is an index change, not a calibrated runtime ratio.

Eligibility requires an explicit series flag, supported status, a classical model, a numerical formula and a primary source URL. There is no practicality filter. Algebraic assumptions, randomness, memory and distributional guarantees are recorded, not turned into confidence percentages. Source checking was performed by an agent; independent human verification remains outstanding.

Scope differs across the five numerical tracks

Reference sizes and the limits of comparability.
TrackWhat the index describes
ECDLPGeneric group-operation estimates at 256-bit subgroup order. Negation changes a constant, not the square-root exponent.
LatticesSelected heuristic sieve bounds at SVP oracle dimension 406, starting in 2008. This omits full LWE reduction, repeated calls, dimensions-for-free and hardware costs. The frontier's refined target sieve dimension 375 is a separate convention.
DecodingSelected full-distance bounds at length 1346, maximizing the exponent over code rates for each algorithm. The hardest rate can change; this is not a single rate-1/2 code or a deployed McEliece/HQC parameter.
FactoringSelected leading asymptotic expressions at a 3072-bit modulus, starting with quadratic sieve in 1982. The incomplete earlier history is excluded. Asymptotic variants are included even if impractical; the index is distinct from NIST's security-strength convention.
Boolean MQSelected bounds for square quadratic systems at 184 variables. Conditional algebraic bounds and general randomized algorithms have different guarantees, stated in the ledger.

Dating and corrected claims

Dates attach to the quoted analysis. A later improved analysis of an old algorithm is not silently backdated. Where the first version has not been established, the ledger identifies the checked source version and the dating uncertainty. Earlier algorithm dates can be recorded separately. Consequently a late point can describe a correction or stronger analysis, rather than a new implementable algorithm.

The original Both–May coefficient 0.0885 is retained only as a corrected claim. The selected corrected coefficient is approximately 0.0951; allowing the obsolete smaller value into a running minimum would preserve the error forever. Earlier worst-case enumeration and later heuristic sieving are not combined into the former 528-bit lattice “loss.” Missing earlier milestones are acknowledged by starting a selected series later, rather than filling the gap with an unsupported flat baseline.

No “years stable,” exponent-change count or security forecast is inferred. Dimensions-for-free provides a subexponential, dimension-dependent saving while leaving the leading linear sieve coefficient unchanged. Memory tradeoffs, implementation work, new proofs and changes in full attack strategy can matter between the plotted steps.

Events without a common cost model

Small-characteristic DLP, pairing-field reassessments and SIDH appear as sourced events. The former finite-field \(2^{70}\) placeholder and SIDH \(2^{42}\) runtime conversion are removed. The approximately 59-bit pairing estimate concerns \(\mathbb F_{2^{4\cdot1223}}\), in the source's modular-multiplication units; the curve's base field \(\mathbb F_{2^{1223}}\) is a different object. Quasi-polynomial time remains super-polynomial. Without a suitable numerical model, inverse size is “not estimated,” never “no finite size.”

Inspect: schema and conventions, generator, full source ledger, regression checks.

What the website adds or changes

The site data builder reads the committed combined export and annual author table. It selects years 1970–2026, rounds annual paper weights to two decimals, retains unrounded full-corpus totals for the summary table, and recomputes named-problem estimates from their labels. It does not harvest sources or run classification models.

The annual paper plot uses a centered three-year mean. At the ends, it averages available years only. Cumulative paper mode sums unsmoothed annual values. Author plots already inherit five-year windows from the estimator. The 2026 observation year is incomplete and is not scaled up to a full calendar year. Cumulative author plots display years from 1980 while retaining earlier contributions in the accumulated values.

The cumulative crossover annotation identifies the most recent year when lattices exceed ECDLP after previously being at or below it. On logarithmic cumulative axes, values below one are drawn at one to permit plotting; this is a display floor, not an additional contribution. The recent annual-rate table averages the five displayed annual estimates for 2022–2026.

The budget slider permits 1×–1000× productivity, $100,000–$500,000 per human year and a hypothetical $1–$100 per million reference tokens. The 1×/10×/100× buttons are examples, not fitted bounds. The problem comparison shows conditional successful-run means, with its flagged AI-literature amounts available in the data table. Ratios that could be mistaken for measured speedups are not printed.

The current budget explorer offers dollar and reference-token views on fixed axes. Price changes the assumed token allowance, \(T=10^6D/p\), while the dollar budget \(D=cH/s\) stays unchanged. Ten times the price buys one tenth as many reference tokens. The HAWK marker in the token view applies the same conversion to its disclosed dollar budget; it is not a reported token count. Using the same price on both sides leaves their ratio unchanged.

The frontier rows in the website builder are curated constants with explanatory metadata. They are not read dynamically from the record-normalization report. Re-running the site builder therefore does not update them automatically after a new record or target model changes. The generated date is a build date, not proof that every source was checked that day.

Inspect: data builder, current chart transformations, current generated payload.

Known defects and choices open to challenge

These issues are part of the interpretation of the published snapshot. Listing them does not assign a numerical correction or prove that they cancel.

A concrete review agenda.
IssueKnown statusWhat would resolve it
Publication share as time shareMay substantially overestimate actual work; no calibrated error bars.Independent time estimates for a diverse sample of projects and contributors.
Author-year aggregate capMaximum-weight guard can fail to cap summed class shares.Restore raw profiles, evaluate affected years and publish a versioned rerun.
Scope/attribution errorsA bounded 45-item keyword review found one clear HAWK/LIP mismatch and several nonce-leakage/other attribution questions.Item-level review of abstracts/full texts, followed by consistent downstream regeneration.
Incomplete leakage metadata39 main-view records explicitly mark leakage; 871 lack the flag.Review missing flags before claiming a comprehensive no-leakage sensitivity.
Source-subset mismatchTen mixed-source records treated differently in old reports and the export.Unify the rule and version the comparison; all-sources headline unaffected.
Unallocated tagsRank/other contribute nothing, LPN contributes half, rsa_variant is unmapped.Decide the intended population before changing attribution weights.
Problem chronology and AI workSome displayed “years” are assigned to AI-assisted publications; dates are unaudited.Define result-specific historical cutoffs and audit provenance and variants.
Cost selectionSuccessful-run costs omit failures; human and AI outputs differ.Complete attempt ledgers and matched deliverables, without assuming cross-field transfer.
Duplicates and identity errorsTitle heuristics and unresolved names can overmerge or oversplit.Trace high-impact records and author matches to stable identifiers.
Computational extrapolationMost records lack human verification; model units and hardware ratios differ.Verify records, make memory/operation units explicit and test alternate models.

The scope audit's 45 keyword-selected records are not 45 errors. Most fit the deliberately broad hint-model or solver scope, and no clear pure timing/power/fault-extraction inclusion emerged from that bounded inspection. The confirmed example is ePrint 2026/2381, “The Lattice Isomorphism Problem with Hints,” allocated 0.75 to structured lattices and 0.25 to plain LWE despite the main rubric's LIP exclusion. The current corpus has not been relabeled. Weighted item counts cannot simply be subtracted from author-year totals, whose denominators and windows are coupled.

Private and classified effort remain unobserved. There is no evidence here that their distribution mirrors public effort, or that agencies compensate most where public effort is lowest. Our comparisons concern documented public work only. The earlier HAWK rediscovery pilot produced no completed mathematical evaluations after safeguard interruptions; it supplies no success-rate calibration. The current scenario method does not require that experiment.

Reproduce, inspect or challenge a number

Start with the pinned repository snapshot for the corpus and effort numbers discussed here. The algorithmic history uses the current repository. The current repository may contain later corrections. The white paper's manifest records input hashes and derived values; it is a provenance ledger, not a scientific validation certificate.

# From the repository root; rebuilds outputs from existing labels.
python3 -I analysis/attack_stability.py
python3 -I pipeline/build_site_data.py
make -C paper

# Separate computational-record report, using its own inputs.
python3 -I analysis/records_normalize.py

The site builder uses Python's standard library. The paper build also requires latexmk and the LaTeX packages declared in its source. It writes intermediate files under tmp/pdfs/whitepaper/ and copies the PDF to output/pdf/ and docs/assets/. These commands do not call a model, spend inference funds, reclassify the corpus or update literature acquisition. Inspect the scripts and dependencies before running the separate record pipeline.

A full acquisition or author-estimator rerun is a different task. It requires raw bibliographic and enrichment inputs omitted from Git, and may encounter changed APIs, database contents and indexing. The committed exports support inspection and figure regeneration; they do not make the entire original acquisition independently reproducible from a clean checkout alone.

Where to begin an item-level challenge.
QuestionEvidence to inspect
Why is this paper counted?combined_included.csv, the original label/reason, rubric, inclusion category and source metadata.
Why does it count toward this class?Problem tags and weights, attribution.json, and any missing or partial mapping.
Why so many research-years?authors.csv, authorships.csv, fte_by_year.csv, career matching, numerator/denominator windows and fallback basis.
What does this AI cost include?ai_results.yaml, erdos_problems.csv, source tables, model/version and success/failure scope.
Why this frontier gap?Record YAML, verification status, target parameters, model formula and omitted overhead.
Why does the chart differ from a report?Snapshot, source filter, year limits, inclusion rule, rounding and chart transformations.

A useful correction names a record or output, identifies the rule or assumption being challenged, and explains whether the problem is factual metadata, classification, attribution, time estimation or interpretation. Those require different repairs. Questions and corrections can be filed in the repository issue tracker.