εCryptanalytic Effort
automated analysis · working draft · data as of 2026-10-09

How hard have we looked?

Measuring public cryptanalytic effort against the assumptions behind standardized public-key encryption

An AI-conducted analysis, overseen by Matthew Green

Abstract

We estimate how much public cryptanalysis each public-key assumption has received, using the literature itself as the instrument. From roughly sixteen thousand candidate papers we classify cryptanalytic work on ECDLP, structured and plain lattices, syndrome decoding and their comparators, then convert papers into estimated research-allocation years through a person-level author map. These publication-based estimates may substantially overstate actual time spent. Lattices have overtaken elliptic curves in this cumulative measure only recently; we express that history as inference budgets under stated productivity assumptions. The study was conducted by AI models: Claude Fable 5.1 and Claude Opus 5.5 built the pipeline, ran the analysis and wrote this page, with contributions from GPT-6 Astra. Matthew Green set the questions, made the key methodological decisions and oversaw the work.

2,877
Corpusitems in the main effort view
~3,240
Peopleestimated distinct author identities
1.7×
Lattices / ECDLPestimated research-allocation years
269
P-256 gapwork beyond the largest public solve

Introduction

Every deployed public-key scheme rests on a conjecture that some problem is hard. We rarely ask how hard anyone has actually tried. This page reports a literature-based estimate of public cryptanalytic effort for each assumption class behind standardized encryption: the elliptic-curve discrete logarithm (ECDLP), structured and plain lattices (MLWE, LWE), and syndrome decoding, with factoring, isogenies and multivariate systems as comparators.All figures count public work only. Classified effort is discussed in §6.

The goal is a defensible order-of-magnitude comparison of the form class X has received about N times the scrutiny of class Y, not a security level. Computational records and model-based work estimates get their own track.

A research-allocation year (called research-FTE in the underlying data) is an estimate based on the share of a researcher's publications devoted to a topic. It is not a recorded year of work.

The literature

We harvested IACR ePrint, DBLP, OpenAlex and zbMATH, screened by keyword, and classified candidates against a written rubric. A second pass reviewed included and borderline items plus a sample of exclusions. Figure 1 shows the result by year of publication; papers that address more than one assumption are split fractionally.

Papers per year
Cryptanalytic research items by assumption class, fractional counts. Annual values are a centered three-year mean; 2026 is incomplete. ePrint / DBLP restricts sources, not correctness or peer-review status. Items without a year or outside the plotted period are omitted from this figure. Click a legend entry to hide a series.

Lattice research exceeds ECDLP in the main cumulative views. The lattice/ECDLP ratio ranges roughly from 1.3 to 2.6 across the study's counting and source assumptions.

These are different measures and sensitivity choices, not a confidence interval. The main paper-count view crosses in 2017; the author-output estimate crosses in 2019.

Visible effort by assumption class. Fractional item counts (include = yes, effort categories); estimated allocation years from the author map. Totals include items missing from the dated chart.

Who does the work

Paper counts treat a two-page note and a decade-long program alike. The author map groups 7,613 authorships into about 3,240 identities, some unresolved, and estimates each person's share of publications devoted to a class over a moving five-year window. Two people with half their output in a class contribute one allocation year.

These estimates may substantially overstate researcher time. Publishing only on a topic does not mean working on it full time. Teaching, administration, other employment, short projects and unrecorded work can all make publication share exceed time share. A minimum assumed publication rate limits the effect of sparse records, but does not measure hours. The alternative career-adjusted view reduces shares by assumed research-time fractions; it is also a proxy, not a correction validated against researchers' time records. Some authors lack career data and use assumed effort per paper instead.

Estimated research allocation per year
Publication-based allocation estimates, not observed staffing. Career-adjusted applies assumed research-time fractions by career stage. Both views retain the per-paper fallback. 2026 is a partial year.

Cumulative effort

In this model, factoring and FF-DLP together still lead on accumulated allocation years, while lattices passed ECDLP in 2019. These curves describe the model's allocation of published output over time.

Cumulative estimated research-years
Cumulative allocation estimates, displayed from 1980; earlier contributions remain in the totals. Hover for values; the marker shows where lattices pass ECDLP in the selected model.

Distance to the frontier

Effort is one axis; demonstrated computation is another. For each target we use a cost model to estimate the work ratio between an attack and a relevant public computational record. We express this ratio in bits:

$$g \;=\; \log_2 W(\text{target}) \;-\; \log_2 W(\text{largest public solve})$$
(1)

A gap of 10 bits means about a thousand times the modeled work; 35 bits means about 35 billion times. Both endpoints use the same model. Only a shared, size-independent cost factor cancels; memory, communication and other costs can change with problem size.

Frontier gap by target
Approximate model-based work gaps. Ranges reflect selected model assumptions, not confidence intervals. Lattice gaps cover one sieve call; the full ML-KEM-512 attack adds roughly nine bits in this model. RSA uses the leading GNFS expression; HQC uses a coarse decoding model.

These gaps are not security levels or forecasts of attack cost. Different families use different operations, memory requirements and model assumptions, so similar numerical gaps need not mean similar practical distances. RSA-2048's roughly 35-bit gap is a GNFS extrapolation, not a measured runtime ratio.

See the normalization method for the models, sources and verification limits.

Attack history and stability

How have attack algorithms changed? The plot below traces selected, source-checked mathematical bounds. Each formula gives a leading-term index: a simplified measure that drops lower-order factors. A fall in this index records an improvement in the selected model. It does not measure security bits lost, and a flat line does not mean that practical attacks stopped improving.

The first view holds a reference size fixed. The second finds the size giving an index of 128 under each selected bound, shown relative to the latest such size. These are illustrative model sizes, not recommended cryptographic parameters. Dates follow the quoted analyses, including later corrections, rather than reconstructing what researchers believed at each moment.

Selected algorithmic bounds over time
A selected history, not an exhaustive best-known frontier. ECDLP uses generic group-operation estimates; lattices use an SVP oracle of dimension 406, not a full ML-KEM attack. Decoding maximizes the full-distance exponent over code rates separately for each algorithm; it is not one fixed-rate code. Factoring and MQ use their stated asymptotic and algebraic assumptions. Levels across these models are not calibrated to the same operation or memory cost. Sources are available in the data table.
Selected model indices. Different starting years reflect source coverage and comparison scope. “Latest plotted bound” is not a claim that no other progress occurred after that date.

An unchanged leading exponent is a narrow kind of stability. Better reduction methods, dimensions-for-free, memory tradeoffs and implementations can improve finite-size attacks without changing that exponent. These histories cannot predict future breaks or convert publication counts into a security guarantee.

Breaks and reassessments with different cost models

Finite-field advances and the SIDH break are retained as sourced events. Their concrete estimates and runtimes concern particular instances and units, so they are not joined to the index curves. Polynomial and quasi-polynomial algorithms are distinct; neither label alone supplies a numerical parameter recommendation.

Selected events. Estimates retain their source's scope; they are not common-unit measurements.

Rules, exclusions and corrections: detailed methodology. Full dated ledger and sources: algorithmic history report. Sources were checked by an agent; independent human verification remains outstanding.

What this does not count

Our literature estimates cover publicly documented work: papers, preprints, theses and published records. Unpublished failures, private research and classified work are missing. We do not have a reliable allocation of that missing effort by assumption class.

If public effort were known, additional classified effort could be represented by an unknown multiplier:

$$H_X^{\text{public+classified}} \;=\; (1 + k_X)\, H_X^{\text{public}}, \qquad k_X \ge 0 \text{ unknown}$$
(2)

We do not know whether this multiplier is similar across classes. Comparisons of the visible literature remain comparisons of public work; extending them to total effort requires additional assumptions. Missing work does not make our publication-based time estimates lower bounds: those estimates may themselves be too high.

Effort in inference dollars

Recent AI systems have produced new results in mathematics and in cryptanalysis at disclosed costs. We do not estimate an exchange rate between human and machine research. Instead we express historical effort \(H_X\) as an equivalent inference budget under an assumed productivity multiplier \(s\):

$$D_X(s) \;=\; \frac{c \cdot H_X}{s}, \qquad c = \$250{,}000 \text{ per research-year}$$
(3)

The multiplier means assumed research productivity per dollar: at 10×, one estimated human research-year corresponds to $25,000 of inference. No multiplier is measured or preferred. These budgets inherit any overestimate in the human-effort proxy: if actual time is half our estimate, the corresponding budgets are also half as large.

Budget explorer
log scale, 1× to 1000×
Hypothetical blended price. Changes token totals, not dollar budgets. Select “reference tokens” to see the chart move.
Scenario budgets using the unadjusted output-share estimates above, with reference tokens at a hypothetical blended price. These are sensitivity cases, not predictions of attack cost. Human supervision, verification, model training and tool compute are outside the inference allowance. The HAWK marker is a disclosed discovery budget, not a calibration of the multiplier. In the token view, that marker converts the same dollars at the selected price; it is not HAWK's reported token usage. Both axes stay fixed as the sliders move.

Read in reverse, a $100,000 inference budget represents 0.4, 4 or 40 assumed human-year equivalents at 1×, 10× or 100×. These are assigned equivalents, not observed labor displaced by AI. The paper separately explores HAWK normalization using paper-based effort on both sides; it does not apply that rate to the author-output estimates shown here.

Recent mathematical results

For selected mathematical problems and HAWK, we compare disclosed AI budgets with a paper-based estimate of effort in the surrounding literature. This second estimator assigns assumed years to each paper, rather than reconstructing author publication shares. The samples include partial results and related problem variants.

Long-standing does not necessarily mean hard-fought. An Erdős problem can remain open for decades because few people have seriously tried to solve that particular question. A low-cost AI solution may therefore reveal an overlooked opportunity, rather than overcome decades of concentrated human effort. Some problems have attracted substantial work; others have not, and we do not know the effort spent on unsuccessful attempts. Papers on a surrounding topic may also devote little time to the exact question. Our literature estimates cannot resolve that distinction. These selected successes should not be treated as evidence that AI can cheaply replace an equally large body of sustained cryptanalytic scrutiny.

Literature-effort scenarios vs. AI budgets
Hollow dots value the literature estimate at $250k per assumed year. Ranges are sums of effort-prior endpoints, not confidence intervals. Filled dots show average disclosed successful-run costs for the Erdős cases and Anthropic's reported HAWK discovery budget. Unsuccessful attempts and other project costs are not included in the Erdős averages. Rows without a budget are literature estimates only. An asterisk marks a literature total containing explicitly AI-assisted work.

The literature totals are not yet audited pre-discovery human labor. The retained corpus includes explicitly AI-assisted papers, marked in the chart, and its historical cutoffs still need review. Existing literature also supplies tools to the AI. Comparing these dollar amounts does not measure productivity gains or the fraction of historical scrutiny reproduced.

Sources: FrontierMath Erdős, Tables 2–3 and Anthropic's HAWK report. The Erdős paper reports more than $220,000 across its wider campaign; the individual successful-run costs shown here are a different accounting boundary. HAWK's roughly $100,000 is API cost for discovery, not the cost of executing a key-recovery attack.

Method and caveats

For the exact rules, formulas, source files and unresolved defects, read the detailed methodology and audit trail.

Corpus. IACR ePrint and DBLP harvests, OpenAlex phrase searches and citation expansion, and zbMATH. The export contains 4,349 included or borderline records; 2,877 are included items in the main effort categories. Recall against twelve selected seed papers' reference lists is approximately 96–100% per class; this does not establish field-wide recall.

Classification. The second model pass reviewed included and borderline items and a sample of exclusions, with access to the first labels. Agreement of \(\kappa = 0.84\)–\(0.89\) describes consistency on the reviewed subset. Exclusion reversal rates of 0.5–1.1% concern sampled first-pass exclusions. Neither check establishes independent validation, classifier recall or a global error rate. Systematic human validation remains outstanding.

Scope. The rubric excludes attacks whose contribution is extracting power, timing or fault leakage, and implementation or protocol flaws unrelated to the mathematical problem. It intentionally includes mathematical attacks given partial information, quantum algorithms, special or weak instances, and faster implementations of mathematical solvers. These broader categories do not all transfer to ordinary deployed parameters. Abstract-based classification can also admit out-of-scope work or misattribute mixed papers; such errors can overstate a class's relevant scrutiny. See the classification rubric.

Time. Output share may substantially exceed time share. Missing career publications, uncertain identities and effort-per-paper assumptions add further error. Career-stage adjustment lowers the principal-class totals by roughly a quarter, but relies on assumed research fractions and often estimated career stages. No defensible numerical error bar is available. Dollar scenarios scale directly with these uncertain totals.

Models. Pipeline, analysis and text by Claude Fable 5.1 and Claude Opus 5.5, with contributions from GPT-6 Astra; smaller Claude models ran the first classification pass. Methodological decisions were made and recorded by the overseeing human.

Data and citation

Read the detailed methodology for the audit trail and the working white paper (PDF) for the synthesis, nominal person-hours, and additional OpenAI and Anthropic comparisons. Its frontier-record analysis remains a separate track, summarized here with the qualifications above.

Every figure on this page is drawn from the generated data. The study repository contains the sources, classification rubric and pipeline/build_site_data.py used to regenerate it.

@misc{bitsofwork2026,
  title        = {How Hard Have We Looked? Measuring Public Cryptanalytic Effort},
  author       = {Green, Matthew},
  note         = {AI-conducted analysis, overseen by Matthew Green. Working draft},
  howpublished = {\url{https://bitsofwork.org}},
  year         = {2026},
}