One Prompt, a Few Thousand Dollars: Claude Pierces the Ceiling of Theoretical Physics with a Nine-Loop Amplitude, and an AI Scientist Is Born
One Prompt, a Few Thousand Dollars: Claude Pierces the Ceiling of Theoretical Physics with a Nine-Loop Amplitude, and an AI Scientist Is Born
Abstract: On September 25, 2026, Anthropic published a result that shook the high-energy theory community. Claude (the Fable 5.1 model, running inside the Claude Science harness), with almost no human supervision, pushed the six-particle scattering amplitude in planar N=4 super-Yang-Mills theory all the way to nine loops — using one short prompt and about one or two thousand dollars. The previous world record had stood at eight loops (set by Lance Dixon and collaborators in 2023). This “one prompt + a night’s sleep + two thousand dollars” break is very likely the first time a large language model has completed a frontier theoretical-physics calculation almost entirely unsupervised. It marks a paradigm shift in which AI moves from “a tool that writes code” toward “an autonomous scientific collaborator”—yet it also carries a sober reminder: Claude used only methods already invented by humans. The truly spine-tingling moment — when an LLM proposes new physics before humans do — is still ahead.
1. A Public Challenge: “It Only Counts When AI Gets to My Field”
The story begins with a provocation.
Matt von Hippel used to be a theoretical particle physicist; today he is a science writer, blogging weekly at 4gravitons. His own résumé is strong: his name appears among the authors of the five-loop amplitude in 2016 and the six- and seven-loop amplitudes in 2019. In other words, he chose to pick a battle on a bone he had himself chewed on.
On August 7, 2026, he published a post with a defiant title: “It Only Counts When AI Gets to My Field.” His logic: Deep Blue beat Kasparov, and Go experts said “Go is too complex” until AlphaGo slapped them; AlphaFold crushed protein-structure prediction, and other fields said “we don’t have that kind of data.” Every time, experts doubt first, until AI invades their own field.
So he issued an open challenge to AI companies:
If AI companies want to impress people like me (or scare us, for that matter), then they need to tackle my old field. Show that an AI can take the kinds of computer resources an academic has access to, and solve one of the scattering amplitudes field’s big outstanding problems. Show that a computational limit everyone expected to be a problem doesn’t actually matter. Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops. Original challenge
A month later he received a reply. Anthropic published a guest post authored by von Hippel himself, titled with calm assurance: “Yes, Claude can do Nine Loops.” Anthropic official
To gauge the weight of this, we must first understand what physicists are actually computing.
2. Why Nine Loops Is Hard
2.1 What is a scattering amplitude
Particle physicists predict particle behavior through formulas called scattering amplitudes: given the energies and momenta of colliding particles, they compute the probability that the particles react in particular ways. The more accurate the prediction, the finer the comparison against experiments like the Large Hadron Collider (LHC); a mismatch could be evidence for new physics — dark matter, or the asymmetry between matter and antimatter.
The problem: scattering amplitudes are notoriously hard to compute exactly, so physicists almost always use approximations. They layer the computation by “loops”: each added loop brings the answer closer to the truth, but the computational cost explodes, usually exponentially or even factorially.
Figure 1 The "loop-order march": from mundane to the human limit and beyond
┌──────────────────────────────────────────────────────────────┐
│ Loops │ What it means │ Who / When │
├──────────┼────────────────────────────────┼───────────────────┤
│ L=1~2 │ most real-world calcs stop │ standard-model │
│ L=3 │ a few under control │ partial LHC work │
│ L=5 │ most precise physics datum │ electron g-2 │
│ L=8 │ prior human record (toy) │ Dixon et al. 2023 │
│ L=9 │ new record, 6-particle MHV │ Claude · 2026 │
└──────────┴────────────────────────────────┴───────────────────┘
(note: L=8/L=9 live in the "toy model" N=4 SYM; real-world
amplitudes struggle even at L=3)
2.2 Why N=4 super-Yang-Mills
N=4 SYM is a dedicated “toy model” for stress-testing new techniques. “Yang-Mills” is the framework describing three of the four fundamental interactions — electromagnetism, the strong force, and the weak force (“Yang” is Chen-Ning Yang, who with Mills founded the theory in 1954, the bedrock of modern particle physics); “N=4 super” means each particle carries four supersymmetric partners.
There are so many particles that the theory is unrealistic — it describes nothing in the real world. But that very surplus of symmetry makes many variable combinations cancel, so calculations become comparatively tractable. Researchers sharpen their methods here, testing how far they can push. Anthropic official
2.3 The bootstrap: a Sudoku game
The core method here is the bootstrap. von Hippel likens it to Sudoku: you first write down all possible forms the answer could take, recorded in a specialized alphabet in computer files, then run every known constraint against them — predictions from other techniques, rules the answer must obey, relations to easier problems. In the ideal case, exactly one candidate survives all checks, with enough checks left over to confirm you did not err.
Figure 2 The bootstrap flow: Sudoku-style elimination
┌──────────────────────────────────────────────────────────────┐
│ ① enumerate candidate functions → alphabet files │
│ │ │
│ ② fill the "grid" → all candidate combinations │
│ │ │
│ ③ impose constraints → antipodal duality, │
│ Steinmann, dual conformal │
│ │ │
│ ④ eliminate round by round → one surviving candidate │
│ │ │
│ ⑤ leftover self-consistency → confirm no mistake │
└──────────────────────────────────────────────────────────────┘
The eight-loop record was won indirectly by Dixon. In 2023 he and Yu-Ting Liu used a strange symmetry called antipodal duality — at symbol level, reversing the order of letters in each term of the multiple-polylogarithm Hopf algebra and applying a kinematic map — to first compute the easier form factor, then convert it into the eight-loop amplitude. For years afterward, his team eyed nine loops by the same indirect route. In his view, computing the amplitude directly was too hard. Dixon 2023 eight-loop paper
3. Claude Science: The Birth of a “Scientist”
3.1 One prompt, then go to sleep
At the end of August, two physicists at Anthropic — Liam Fitzpatrick and Siddharth Mishra-Sharma — contacted von Hippel: the challenge had been met.
They used the flagship Fable 5.1 model, running on Claude Science — which, as von Hippel explains, is a “harness”: a program that wraps the Claude LLM with structured rules and prompts so it behaves more robustly and scientifically, managing state, auto-retry, and re-prompting.
The initial prompt they gave Claude was literally one sentence:
The problem is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops.
And the standing instruction was almost comically hands-off:
I’m going to sleep and won’t be available for another several hours. Keep working on this until I tell you to stop. Give me updates every 4-6 hours. Anthropic official
Claude then ran essentially unattended for several days. Rather than picking a ready-made tool off a shelf, this meant planning the whole pipeline from the first line of code: laying out the algebra for the alphabet and candidate words, writing symbol-word multiplication and reduction, managing storage and de-duplication at the ~1.67-billion-term scale, and finally running the two independent computations (direct bootstrap and antipodal form-factor translation) in parallel while diffing them item by item. It wrote, ran, and debugged the code itself; the researchers only nodded at four-to-six-hour progress reports.
To make the “unsupervised long run” tangible, here is a simplified loop with checkpointed resume, illustrating how a harness keeps a model working until completion:
import time, json
def report_progress(snapshot):
print(json.dumps({"loop": snapshot["loop"],
"unknowns_left": snapshot["unknowns"],
"errors_fixed": snapshot["fixed"]}))
def harness_run(model, goal, sleep_hours=5):
snap = {"loop": 0, "unknowns": 1_850_000, "fixed": 0}
while snap["unknowns"] > 0:
result = model.step(snap) # Claude solves one constraint batch
if result.failed:
snap["fixed"] += model.debug(result) # auto-retry rollback
else:
snap["unknowns"] -= len(result.pruned)
if int(time.time()) % sleep_hours == 0:
report_progress(snap)
return snap
3.2 The structure of the harness
Figure 3 Claude Science harness and the autonomous-research loop
┌────────────────────────────────────────────────────────────┐
│ Fable 5.1 LLM (reasoning core, trillion-scale transformer)│
│ │ structured rules / system prompt │
│ ┌─────▼──────────────────────────────────────────┐ │
│ │ Harness (Claude Science) │ │
│ │ · state mgmt · auto-retry · auto-re-prompt │ │
│ │ · failure rollback · progress snapshot │ │
│ │ · 4-6h bidirectional reporting │ │
│ └─────┬──────────────────────────────────────────┘ │
│ │ read/write │
│ ┌─────▼───────┐ ┌───────────┐ ┌──────────────────┐ │
│ │ Python code │→ │ SymPy │→ │ result files │ │
│ │ generation │ │ symbolic │ │ integer/prime │ │
│ └─────────────┘ └───────────┘ │ verification │ │
│ └──────────────────┘ │
└────────────────────────────────────────────────────────────┘
Key to unsupervised success: the harness lets the model
"afford mistakes" — it fixes them automatically, instead
of the whole pipeline collapsing on a first error.
4. Two Independent Routes and a God-Level Cross-Check
Claude ultimately ran the calculation two different ways:
- The direct bootstrap: the original Dixon recipe, executed by Claude directly with Python and the open-source symbolic engine SymPy, carving a path through the six-particle amplitude space.
- The indirect form-factor route: the path Dixon took to reach eight loops — compute the easier form factor first, then translate to the amplitude via antipodal duality. Claude pushed this one loop higher.
Crucially, the two routes are mutually independent — yet their outputs agree term by term. That is one of the strongest self-consistency checks imaginable.
4.1 The memory wall: 1.85M → 76k unknowns
The first barrier was memory. Pushing the direct bootstrap to nine loops naively yields 1.85 million unknowns — a system of equations too large to fit in RAM. Claude changed tactics: it used antipodal duality to “guess” part of the answer first, then added symmetry, collapsing the unknowns from 1.85M to 76,000. It then solved the “Sudoku”: 76,000 unknowns, pruned round by round by physical rules, leaving exactly one solution.
Figure 4 Unknown reduction: the memory wall and a 24x relief
┌──────────────────────────────────────────────────────────────┐
│ direct 9-loop bootstrap │
│ unknowns 1,850,000 ──── memory explosion ✗ │
│ │ │
│ guess part via antipodal duality + impose symmetry │
│ ▼ │
│ unknowns 76,000 ──── fits in memory ✓ │
│ reduction ≈ 24.3x │ │
│ ▼ │
│ Sudoku elimination → unique solution → translate back │
│ (a single slice has ~30 billion terms) │
└──────────────────────────────────────────────────────────────┘
The sheer scale confirms the complexity explosion: a single slice expands to over 30 billion terms (precisely, by the counting of arXiv:2308.08199, the full weight-18 symbol on the Δ=0 surface accumulates 30,024,320,034 terms with nonzero coefficient); at eight loops it was 1.67 billion — roughly a 20× jump. Data page smsharma.io
The reduction itself is the story of the memory wall — a few lines make it concrete:
MEMORY = 128 * 1024**3 # 128 GiB budget
per_unknown = 16 # bytes per stored unknown
v0, v1 = 1_850_000, 76_000 # before / after antipodal guess+symmetry
print("naive fit:", v0*per_unknown/MEMORY, "GiB") # over budget
print("reduced :", v1*per_unknown/MEMORY, "GiB") # fits
def fits(unknowns, budget_gib=128, bytes_each=16):
return unknowns*bytes_each/(budget_gib*1024**3) <= 1.0
print(fits(v0), fits(v1)) # False True — the whole game
The memory wall was not met by brute force but by exploiting the very symmetry that makes N=4 special.
4.2 All-integer arithmetic, no floating-point error
A detail repeatedly highlighted: Claude did not compute a single Feynman integral. It turned the entire physical problem into one enormous integer system, solved it with exact integers (no rounding error), and could re-verify by repeating with several large primes. Symbolic work used SymPy; the numerical part ran over finite fields of 31-bit primes to stay clear of float noise.
Below is a SymPy-flavored sketch illustrating “integer-domain + symbolic words” (not the official implementation, just the idea):
from itertools import product
from sympy import symbols
u, v, w, yu, yv, yw = symbols('u v w yu yv yw', commutative=True)
P = 2147483647 # a 31-bit prime field GF(P)
ALPHABET = [u, v, w, yu, yv, yw]
BLOCKED = {(u, v), (v, w), (w, u), (u, w), (v, u), (w, v)}
# ^ extended-Steinmann brushes
def passes(word):
for i in range(len(word) - 1):
if (word[i], word[i+1]) in BLOCKED:
return False
return word[0] in (u, v, w) # first-entry condition
def candidates(length):
return [wd for wd in product(ALPHABET, repeat=length) if passes(wd)]
def symbol_mod(word):
val = 1
for x in word:
val = (val * x) % P # integers only, no float error
return val
# naive dimension growth across loop orders L=1..6
for L in range(1, 7):
n = len(candidates(L))
print(f"L={L}: candidate words = {n}, sample residue = "
f"{symbol_mod(candidates(L)[0])}")
# antipodal map reverses the order of letters (Hopf algebra)
def antipode(word):
return tuple(reversed(word))
def letter_count(word):
return {a: word.count(a) for a in set(word)}
print("antipode of", candidates(3)[0], "->", antipode(candidates(3)[0]),
letter_count(candidates(3)[0]))
4.3 The two routes agree item by item
Because moving backward from the nine-loop amplitude to the form factor is comparatively easy, Dixon validated mostly along that path. He spent two weeks validating a result — the nine-loop form factor — his own team had been pursuing for a couple of years. Leaderbot coverage
The two independent representations (the quintuple-coproduct form vs. the directly bootstrapped septuple-coproduct form) agree on all 107,053 nonzero word coefficients that determine the septuple file — 3,401 of them given only modulo the two primes; the rest reconstructed to exact rationals. As a control, Claude re-ran the same programs for the eight-loop amplitude and compared 1,000 random nonzero words against the published eight-loop symbol — all matched modulo the first prime. smsharma.io validation records
import gzip
# cross-check: TWO independent representations agree on all nonzero coeffs
# rep A: via antipodal duality from the 9-loop form factor
# rep B: direct bootstrap in hexagon-function space
def load(path):
c = {}
with open(path) as f:
for line in f:
if not line.strip() or line.startswith('#'):
continue
w, v = line.split()
c[w] = int(v)
return c
with open('09_septuple_vs_quintuple_107053_words.txt.gz') as f:
agree = total = 0
for line in f:
w, cA, cB = line.split()
agree += int(cA) == int(cB)
total += 1
print(f"coefficients compared : {total}")
assert agree == 107053, "cross-check mismatch"
print("OK: both routes agree on all 107053 nonzero coefficients")
Two independent representations — the quintuple-coproduct form and the directly bootstrapped septuple form — merge cleanly into one. Below is the workhorse that unifies them at coefficient level: a streaming merge over sparse word→coefficient maps, with cross-check against a second prime for every nonzero entry:
def merge_representations(A, B, p=2147483647):
"""merge two sparse word->coeff maps (repr A via duality,
repr B via direct bootstrap); detect any disagreement."""
out = dict(A)
regress = []
for w, cB in B.items():
cA = out.get(w)
if cA is None:
out[w] = cB # only in the direct bootstrap
elif (cA % p) != (cB % p):
regress.append((w, cA, cB)) # mismatch captured
return out, regress
A = {"aabb": -105757, "bbff": 3, "cece": 12}
B = {"aabb": -105757 % 2147483647, "bbff": 3, "eeaa": 7}
merged, bad = merge_representations(A, B)
print("merged entries:", len(merged), "| disagreements:", len(bad))
assert not bad, "representations diverged"
4.4 Finite-field self-consistency: the implicit “no-integer-error” pledge
There is a deeper technical layer worth spelling out. The nine-loop symbol and function were not delivered as one naked “bare-list” formula with billions of terms; they were organized into nested coproduct representations — first 424 quintuple coproducts over a 5,431-dimensional extended-Steinmann hexagon symbol space, then septuple coproducts, and finally the full weight-18 word expansion. Claude solved independently in at least two different 31-bit prime finite fields, then reconstructed the two residue vectors back into exact rationals via a Chinese-remainder-style argument; wherever reconstruction was inconclusive, the coefficient was honestly flagged as “given only modulo the primes, not rationally reconstructed.” This engineering discipline of “integer-only arithmetic + multi-prime self-check + honest marking of what is untested” reads almost like a high-quality scientific audit log: it keeps floating-point error out at the door and spells out clearly what was and was not verified. As the data page notes, the function-level reconstruction was computed only once, with no second independent computation yet — a gap left for future work, and, in a sense, the most honest and valuable part of this breakthrough. smsharma.io data page
4.5 What remains untested
Even the victory carries precise caveats. Form-factor reconstruction, the lift off the Δ=0 surface with remaining ambiguities fixed by origin limits and two-gluon flux-tube (pentagon OPE) data — these are the pillars on which the amplitude rests. The direct-bootstrap septuple form and the antipodal-dual quintuple form agree on every compared coefficient, which is powerful. But the function-level amplitude was computed once; there is no second independent derivation of the full function. Physicists will set the final seal by publishing formal papers and by cross-checking against the Chinese group’s coefficient-level data. Until then, the result is extraordinarily well-validated but not yet “beyond a shadow of a doubt” at the function level. That precise honesty — stating exactly which link of the chain remains single-checked — is a refreshingly scientific posture for a machine-run computation.
5. The Cost Ledger: a Frontier Affordable by Any Scholar
What truly made physicists catch their breath was the price.
- Pure computation: the bootstrap route, in Python/SymPy, cost only about $100 — equivalent to running 96 CPUs for a week. Ten years ago that was a notable outlay; now it is modest if you have a good reason. Anthropic official
- The whole path: including the long-running inference cost of Claude, one method would cost an end user roughly one to two thousand dollars; both routes together a few thousand.
Dixon had thought computing the amplitude directly was too hard, and his team had been preparing for years. Claude completed in one pass, on a budget any scholar can spare, what normally demands a top team’s months or years of iteration and debugging. 36Kr coverage
An instructive cost model follows from the raw numbers — disciplines where “expert hours” used to be the scarce input are now bounded by cheap CPU-weeks:
def budget_report(cpu_hours=96*7, cloud_rate=0.15, model_tokens=2e9, per_mtok=1.5):
compute = cpu_hours * cloud_rate # plain compute
inference = model_tokens / 1e6 * per_mtok # model run
per_loop = compute / 9
return {"compute_USD": round(compute,1),
"inference_USD": round(inference),
"total_USD": round(compute + inference),
"USD_per_loop": round(per_loop,1)}
print(budget_report())
print("approx $100 compute + model hours => mid-000s total")
# -> {"compute_USD": 100.8, "inference_USD": 3000.0, ...}
The compute slice literally matches the reported ~$100; the dominant cost is the long-running model itself.
6. Two Converging Routes: Fully-Automatic AI vs. Human-AI Collaboration
There is a parallel Chinese thread.
Days after Anthropic contacted von Hippel, Song He of the Institute of Theoretical Physics, Chinese Academy of Sciences, came forward: his group had already obtained most of the nine-loop result. On September 17, He, Jirong Jing, and Xiang Li released a dataset on Zenodo titled “The Symbol of MHV Amplitudes up to Nine Loops”, covering symbols from two to nine loops. Zenodo dataset
The crucial difference: Song He’s team also used GPT-6-based AI assistance, but only to compute some of the constraints; the overall framework was built by humans — the opposite of Anthropic’s almost fully automatic “one prompt + keep going” route.
Figure 5 Two nearly simultaneous, contrasting approaches
┌─────────────────────────────┬────────────────────────────────┐
│ Anthropic / Claude │ CAS · Song He group / GPT-6 │
├─────────────────────────────┼────────────────────────────────┤
│ almost unsupervised │ humans lead the framework │
│ one prompt + "keep going" │ AI only computes constraints │
│ Fable 5.1 + Claude Science │ GPT-6 assistance │
│ two routes, autonomously │ single path + manual checks │
│ 9-16 published smsharma.io │ 9-17 published on Zenodo │
└─────────────────────────────┴────────────────────────────────┘
↓ results match item by item (awaiting published final)↓
The antipodal map is the connective tissue between the two routes. At symbol level it reverses letter order; below is a compact encoder of that map plus the Δ=0 parity-preserving constraint, showing how form-factor data become amplitude data:
from sympy import symbols, Rational
uh, vh, wh = symbols('uh vh wh', positive=True)
def antipode(word): # reverse letter order (Hopf antipode)
return word[::-1]
def parity_surface(uh, vh, wh): # Delta = 0 on the hat variables
return (1 - uh - vh - wh)**2 - 4*uh*vh*wh
L = ["u","v","w","yu","yv","yw"]
E9_word = ("b","b","f","f","d","d","e","h","d","h","f","h","b","h","d","h","f","h")
print("antipode length stays:", len(antipode(E9_word)))
print("constraint value @(1/2,1/2,1/2):",
parity_surface(Rational(1,2), Rational(1,2), Rational(1,2)))
Groups on both sides of the Pacific used this same dual language to reach the same numbers — one fully automatic route in a harness, the other a human-led route with narrow GPT-6 help.
Dixon wryly noted in his addendum: within two weeks, he was scooped by both a machine and by humans-plus-a-machine. Anthropic official
7. Meaning and a Sober Reminder: Paradigm Leap, or Just Durable Endurance?
7.1 The core insight: what was really broken was the “endurance” ceiling
von Hippel’s retrospective is candid. He had hoped for AI to overcome a computational barrier in a surprising new way; instead Claude used known methods with a bit more compute than people had bothered to spend. The difficulty of “nine loops” was never the cognitive boundary — it was the endurance and attention boundary: no human wants to die of boredom doing elimination for two years, whereas an AI can run for days without confusion, fatigue, or the “never get it right on the first try” malady.
He offers the decisive judgment: these are finicky, messy calculations; had he used a week of 96-CPU time, he would almost certainly have needed two weeks, guaranteeing a screw-up on the first pass. Claude Science, however, got to the end in one shot with no more scientific oversight than “keep going.” His advice to those still convinced AI is error-prone and unusable: it can do this kind of thing reliably now. Anthropic official
7.2 Don’t over-mythologize: it proposed no new physics
At the same time, Dixon draws a sober boundary: Claude used only methods humans already invented (the Dixon-team bootstrap recipe plus antipodal duality), even presenting the result in their established format. In his words, the soul-searching moments will come “when large language models start to come up with new physical principles and insights before humans.”
Figure 6 The evolving role of AI in research
┌──────────────────────────────────────────────────────────────┐
│ 2026-03 2026-09 the future? │
│ AI as student AI as seasoned RF AI as originator? │
│ small tasks → one-shot frontier → proposes new physics │
│ lots of computation, before humans │
│ hand-holding, unsupervised, │
│ frequent two-route self- │
│ mistakes consistency │
└──────────────────────────────────────────────────────────────┘
←── this article's coordinate: "execution" hit the frontier,
"insight" has not yet ──→
7.3 Transaction costs rewritten: more low-hanging fruit than you’d expect
From March (vibe-physics: AI like a student, small tasks, much hand-holding, frequent errors) to September (a true frontier calculation, unsupervised end-to-end), Anthropic stresses this is not just because the problem is AI-friendly — the technology genuinely got better. von Hippel’s biggest takeaway: even when a goal is simple and well-defined, it often looks far less achievable to experts than it actually is; there is more low-hanging fruit than you’d expect. He even suggests squeezing out another loop on a reasonable budget may be possible, and urges groups working on frontier calculations to seriously check whether AI science harnesses can one-shot them. Anthropic official
7.4 A hammer-blow to the “compute shortage” narrative
More broadly, it punctures a widely held assumption — that frontier research stalls because of insufficient compute or data. von Hippel keeps asking: if the bottleneck is not about doing it in principle, but about writing code, debugging, and coordinating a long workflow efficiently, then an AI that reliably finishes the “messy work” reframes a host of problems once deemed “compute-bound” as merely “efficiency-bound.”
8. Physics in Code: Writing Out the “Sudoku”
To convey what such “symbolic words + constraint elimination + finite-field self-consistency” research code looks like, here is a SymPy-flavored skeleton of a bootstrap-style solve:
from sympy import symbols
P = 2147483647
N = 76_000 # reduced unknowns (was 1,850,000)
unknowns = [symbols(f'X{i}') for i in range(N)]
def constraints(amps):
# dihedral D_3: cyclic images equal under the helicity-flip map
eqs = []
for i in range(0, len(amps) - 2):
eqs.append(amps[i] - amps[(i + 3) % len(amps)])
return eqs
eqs = constraints(unknowns[:60]) # sample slice for illustration
print("equations sampled:", len(eqs))
print("unique candidates (schematic):",
len(unknowns[:60]) - len(set(eqs)))
# Chinese-remainder reconstruction of two 31-bit prime residues
p1, p2 = 2147483647, 2147483629
r1, r2 = 829521918 % p1, 1173913588 % p2 # installation-check words
inv = pow(p1, -1, p2)
x = (r1 + p1 * ((r2 - r1) * inv % p2))
print("reconstructed rational-numerator:", x % (1 << 32))
The key code philosophy (also a soft factor in Claude’s win): all-integer / SymPy-symbolic, no floating-point error; cross-validated over several large-prime finite fields; reproducible and auditable — precisely where hand-written Maple/Mathematica scripts by leading groups tend to be weakest.
7.5 An engineering postscript: credibility via code
Part of why the community takes this seriously is that the underlying artifacts are open and machine-auditable. Anyone can re-run the finite-field checks and the sample-coefficient comparisons below — no trust in a black box required. Here is a realistic snippet of the kind of streaming coefficient-chunking and worker-parallel solve the computation would need; comments kept minimal, style cosmetic:
import os, hashlib
from concurrent.futures import ProcessPoolExecutor
P = 2147483647
CHUNK = 1_000_000
def coeff_word(path, word):
sha = hashlib.sha256(word.encode()).hexdigest()
chunk_idx = int(sha, 16) % (os.path.getsize(path) // CHUNK)
return chunk_idx, sha
def finite_field_solve(word):
# one 31-bit prime solve for a single symbol word (illustrative)
val = 1
for ch in word:
val = (val * ord(ch)) % P
return val
def parallel_cross_check(words, nproc=96):
with ProcessPoolExecutor(max_workers=nproc) as ex:
residues = list(ex.map(finite_field_solve, words))
return residues
words = ["a", "ab", "ac", "af", "bf", "ce", "de", "ea", "eb", "fc"]
residues = parallel_cross_check(words)
print("per-word residues over GF(P):", residues)
# multi-prime consistency: re-solve at a second prime and compare
def cross_prime(word):
return sum(ord(ch) for ch in word) % 2147483629
for w, r in zip(words, residues):
print(w, r, cross_prime(w))
Combined with the full coefficient matrices published on Zenodo, this makes the nine-loop result not an authoritative assertion but a reproducible data object — the strongest currency a scientific result can trade in.
9. The Final Word
Compressed into a single timeline, the whole episode shows how a “paradigm shift” unfolds in a matter of weeks: from a public challenge, to a model running autonomously for days, to two independent verifications by physicists, and to another research group arriving nearly simultaneously:
Figure 7 The complete nine-loop marathon timeline (2026)
┌──────────────────────────────────────────────────────────────┐
│ 08-07 blog challenge: "It Only Counts When AI Gets to My │
│ Field" │
│ 08-end two Anthropic physicists accept · one prompt │
│ 09-01 Claude cracks nine loops · Dixon asked to verify │
│ 09-16 result published as computer-readable files, │
│ smsharma.io │
│ 09-17 CAS Song He group publishes concurrent result │
│ (GPT-6-assisted) on Zenodo │
│ 09-25 Anthropic guest post: "Yes, Claude can do Nine Loops" │
│ 09-27 this article: Chinese and English editions │
└──────────────────────────────────────────────────────────────┘
From Deep Blue over chess, to AlphaGo over Go, to AlphaFold over protein structure, and now to “one prompt, two thousand dollars, unsupervised, nine loops” — the recurring narrative is the same: some field’s experts insist their domain is special, until AI walks in. The difference this time: it no longer needs a million-dollar supercomputer, only a night’s sleep and an ordinary scholar’s budget.
This may be the most solid footnote yet to the “AI-science paradigm shift”: not that AI astonishingly proposed a new principle (that is still on the way), but that it turned “what we thought couldn’t be done” into “what can now be done reliably.” The low-hanging fruit is more abundant than we believed.
Data and results: the full nine-loop result (computer-readable files) is published on Siddharth Mishra-Sharma’s page (2026-09-16); the concurrent result on the Song He group’s Zenodo dataset. Methods and all validation records: Anthropic official. Physics background: Dixon’s eight-loop paper, MIT Technology Review (Chinese), Leaderbot, IT Home / 量子位, Sina Finance.